Friday, November 26, 2010

The Curious Case of Brooks Robinson's Batting Runs (rWAR)

Colin Wyers of Baseball Prospectus pointed this out to me, and neither he nor I have an explanation for it. Rally's WAR estimates have become the most widely used on the internet, especially since they are available at Baseball-Reference. However, some of the batting runs figures don't make a whole lot of sense, and the specific player that Colin brought to my attention was Brooks Robinson.

Rally lists Robinson with a career total of 20 batting runs (above average). That figure does not include baserunning (0 runs) or GDP runs (-35 runs) or reached on error runs (-2 runs), so I will omit those areas of the game from my estimates which follow. The 20 batting runs seems awfully low. My own crude ERP-based estimate is 154 batting runs with park adjustment, 113 without (I estimate Robinson's career season-weighted PF to be .97, meaning he played in moderate pitcher's parks on average). Colin's estimate is 84 batting runs without park adjustment. Pete Palmer's estimate from the 2005 ESPN Baseball Encyclopedia (which does include base stealing) is 53 batting runs. Wyers used OPS+ to generate a crude estimate of +57.

When I was initially discussing this on Twitter, it completely slipped my mind that my figures were comparing Robinson to all league hitters (including pitchers for 1955-1972). Palmer's estimate and OPS+ exclude pitchers from league totals, and they are the closest to Rally's. Still, a thirty-run difference is still fairly large when dealing with offense.

In order to understand why we see discrepancies, it makes sense to attempt to replicate Rally's approach. His explanation of batting runs allows us to get a sense of his process:

Bat runs - This is park adjusted linear weights batting runs, using customized weights at the team level to ensure that total runs credited to players will equal the actual runs scored for that team.

While it is not specified in the quoted entry, Rally has explained elsewhere that he uses Base Runs to generate the linear weights. Since the weights are set so as to ensure that team BsR is equal to actual team runs, I'm going to assume that he's achieving this through the use of a custom B multiplier for each team-season.

I attempted to mimic this through use of a BsR formula that only considered the basic batting events--singles, doubles, triples, homers, walks, and at bats. This formula is far from the most accurate BsR equation ever devised, but it should perform well enough in this role:

A = H + W - HR
B = 2TB - H - 4HR + .05W
C = AB - H
D = HR

Using this equation, I calculated the B multiplier needed for each Orioles season to make BsR = actual runs scored. Then I calculated the corresponding intrinsic weights for each team-season, and used these to estimate Robinson's runs created. From there, I estimated Batting Runs by taking RC - Avg(RC/Out)*Robinson outs.

Using this approach, I estimate that Robinson contributed 107 batting runs (without accounting for pitchers and without a park adjustment).

In order to better mimic Rally's approach, I needed to remove pitcher hitting from the league total. To do this, I used the BsR formula to estimate intrinsic linear weights for each league-season, then figured the league RC/O for non-pitchers (I used a spreadsheet published by Terpsfan101 to get the non-pitcher totals). Using those figures as the baseline for Robinson, I got an estimate of 30 batting runs, which isn't that far off of Rally's. However, when a park adjustment is applied, it shoots back up to 71 runs, which is much closer to the Palmer and Wyers estimates.

More concerning was another curiosity that Wyers noted--the 1969 season. Robinson's Orioles are credited with a team total of 40 batting runs. However, they average 4.81 R/G in a league with an average of 4.09 runs, which means that they scored about 117 runs more than average. That's nearly an 80 run discrepancy!

It gets even more confusing when one looks at the league total listed at Baseball-Reference for the 1969 AL--a whopping -685 batting runs. I have no idea whether this is a problem with B-R's implementation of Rally's method or something else, but it obviously is an error of some sort.

What is different between Rally's figures and my attempts to replicate them? Obviously, if we knew for sure this exercise wouldn't be necessary, but it's safe to assume that:

1. Rally is using a different (and probably better) BsR equation than I am
2. His park factors and mine are probably similar, but surely there are differences
3. Rally may be incorporating some additional categories that I've ignored (intentional walks and sacrifices)

However, the likelihood that all of those differences work against Robinson and account for the difference is not that great (not to mention that the 1969 league figures are illogical). I feel a little guilty posting this without first consulting Rally about it, but I did not have his contact information. He's a good sabermetrician and it is quite possible that I am missing something here--but I do think there is enough smoke to warrant some further explanation.

Monday, November 22, 2010

Meanderings

What follows is a very lightweight post, even for one of this nature.

* I have a half-written post somewhere about the generation gap in sabermetrics between people who got into the discipline prior to the explosion of online sources and those who started at some time after that. I've never finished it or posted it because it's not about baseball--it's about sabermetrics, and because one could easily read it as self-aggrandizing (and perhaps even as a sign of old fogeyism). But the themes have manifested themselves a little bit in the reaction to Felix Hernandez winning the Cy Young.

I'm not crazy about looking at the BBWAA votes for an award as any kind of triumph or defeat for sabermetrics, but if you are inclined to view it in those terms, it's tough to see how Hernandez' win was anything but a victory for the discipline. The win craze in Cy Young voting may have reached its zenith after the Stone/Vuckovich/Hoyt selections stopped in the early 1980s, but it never fully died, not with Jack McDowell in 1993 or John Smoltz in 1996 or Bartolo Colon in 2005. The shiniest W-L may not have been the strong Cy indicator it once was, but a good W-L record was still necessary to get a seat at the BBWAA table (provided the pitcher in question was a starter). It was unprecedented that a 13-12 pitcher would get serious consideration.

It's absolutely true that one didn't need FIP or xFIP or SIERA to make a case for Felix Hernandez; ERA, innings pitched, and strikeouts, which have been kept for the last century, were sufficient to make one consider that Hernandez might have been the league's outstanding hurler. Still, it should not be forgotten that the notion that ERA and strikeouts and the like were useful indicators is one embraced by sabermetrics, that had many less adherents pre-James than it did in 1990, and many less adherents in 1990 than it did in...well, you get the idea.

But for certain members of the community (largely peripheral members, i.e. not the people authoring sabermetric blogs or engaging in their own research), generally those that fall into what could be called (uncharitably to the site) the "Fangraphs generation" of saberites, the notion that actual runs allowed is an acceptable tool by which to evaluate starting pitchers is foreign, as foreign as the notion that W-L was the key evidence was to my generation of sabermetricians.

* Any skirmishes about the baseball awards are a garden party compared to the battle being waged over Horse of the Year between Zenyatta and Blame. I would vote for the latter without a moment's hesitation, and I've yet to see a coherent argument for Zenyatta that is based solely on her 2010 performance. The Zenyatta crowd talks about her "transcending racing" (it's not a popularity contest), or about how she should have won in 2008 or 2009 (arguable, but wrong I believe, and irrelevant to a 2010 award), or about her accomplishments in 2008 and 2009 (beyond irrelevant). Blame ran a more ambitious campaign, beat better horses more times, beat Zenyatta head-to-head, had better speed figures, won more money, and ran exclusively on dirt and at classic distances.

Hernandez/Sabathia is actually not a bad comparison--Sabathia pitched well and wouldn't hardly have been the worst selection in the award's history--but outside of W-L record, it was hard to find an area in which he had Hernandez beat. Outside of the fact that she's Zenyatta, it's hard to find an area where she had Blame beat. To the same degree that I was reasonably confident that Felix Hernandez was the best AL pitcher in 2010, I'm reasonably confident that Blame was the best North American thoroughbred of 2010.

* It now looks as if the expanded playoff format is an unstoppable train. Writing on the idea in an earlier post, I said "In this case, not only do I consider the idea stupid, but it would seriously dampen my own enthusiasm for the playoffs."

Reading it back, I realize that was an overreaction. I don't like the idea of an extra wildcard team any more today than I did then, but I do realize that the likelihood of my enthusiasm for the playoffs being dampened is next to zero. If anything, I'll probably be happy to have a few extra games to watch. The allure of the game is too strong, and to make bold statements about my own ability to resist is self-flattery. I'll object with my head, but I'll tune in and I'll like watching the games if not agreeing with their existence--and so will others, and everyone will make money.

Also, it is worth noting that even with ten playoff teams, MLB will still have the lowest proportion of playoff teams among the big four US leagues.

Tuesday, November 16, 2010

IBA Ballot: MVP

I don't see any slam dunk choice for the AL MVP. My initial RAR numbers have Miguel Cabrera at 74, Jose Bautista 71, Josh Hamilton 68, and Robinson Cano 64. Adding in a crude fielding estimate ((UZR + Dewan's RS)/4) puts Hamilton in the lead at 72, followed by Cabrera 70, Bautista 70, Cano 66, and Longoria 62. Hamilton is also hurt by the fact that the initial RAR considers him a left fielder, but he actually played 22% of his innings in center. Refiguring his position adjustment to take this into account, his offense-only RAR is bumped up by a run, leaving him at 73 total.

It also stands to reason that Hamilton contributed as much or more on the bases than his competitors--BP's EqBRR less stolen base runs (steals are already accounted for in my RC formula) has Hamilton +2, Cabrera 0, Cano +1, Bautista -1, and Longoria +3, and thus only increases Hamilton's insignificant edge. It's not a factor that I consider, but Hamilton will almost certainly win the BBWAA award as he played for a playoff team and Cabrera did not.

There's one player left to consider before handing the award to Hamilton--Felix Hernandez. Hernandez' 76 RAR is definitely comparable to Hamilton's grand total of 75 RAR. However, Hernandez' peripherals are not quite as brilliant as his actual runs allowed, and while I have no qualms about choosing a pitcher as MVP, I like it to be a somewhat clear choice. Since the one run difference in RAR is meaningless and the evidence suggests that Hernandez is getting credit for a fair/favorable runs allowed rate, I can't justify going with him.

The bottom of the ballot is just a matter of mixing in the top starting pitchers with the position players, for whom I see little reason to deviate from RAR ranking. The exception is Paul Konerko who is at 55 RAR but frowned upon by the fielding metrics (-8) and is in front of a bunch of guys for whom I think most people would agree bring a lot more to the table in every area except batting (Adrian Beltre, Joe Mauer, Shin-Soo Choo, Carl Crawford). I would love to be able to justify getting Choo onto my ballot, but Carl Crawford ranks as his equal at the plate and adds more on the field and the basepaths:

1) LF Josh Hamilton, TEX
2) 1B Miguel Cabrera, DET
3) SP Felix Hernandez, SEA
4) RF Jose Bautista, TOR
5) 2B Robinson Cano, NYA
6) 3B Evan Longoria, TB
7) 3B Adrian Beltre, BOS
8) SP CC Sabathia, NYA
9) SP Jered Weaver, LAA
10) LF Carl Crawford, TB

The battle for top position player in the National League can be fairly safely restricted to three first baseman: Albert Pujols (82 RAR), Joey Votto (71), and Adrian Gonzalez (69). Next on the RAR list is Matt Holliday (61). Pujols has a sizeable lead over Votto in my RAR figures, one that may surprise a lot of readers at first glance, and even I was surprised at the margin.

Looking at their unadjusted batting lines, Votto (.324/.420/.600, 8.8 RG) appears to have the slight offensive edge over Pujols (.312/.414/.596, 8.6 RG). However, Pujols still has a four-run cushion in RAR thanks to an extra nine games played and 52 PA. When park is taken into account, Votto (.319/.414/.591, 8.6) and Pujols (.317/.421/.605, 8.9) essentially exchange raw stat lines with one another.

Consider that over the last five seasons, St. Louis's average RPG is 8.8 at home and 9.4 on the road. Cincinnati's split is 9.6/9.1. The parks have played as close to mirror images of one another. Of course park factors can't capture all of the potential influences on those figures--team construction, year-to-year weather fluctuations, chance, etc.--but I don't think it's outlandish to suggest, as my park factors do, that the overall run environment in which Cincinnati plays its schedule is 6% higher than that of St. Louis.

Maybe you don't trust the park adjustment. Maybe you'd prefer to look at each player's performance in the actual run context of his team in 2010, rather than the idealized league average context offered by park adjustments. There are drawbacks to such an approach, most notably that it assumes that each team is equally strong offensively and defensively, but there's an argument to be made that it captures value more effectively than does the neutralization approach. (Bill James made this argument using a fictional Jim Rice as an example in the original Historical Baseball Abstract, and it's something that I intend to ruminate on at some point).

Cincinnati games saw an average of 9.1 runs in 2010 (or 4.55 per team); St. Louis 8.5 (4.25); and throwing in San Diego for good measure, 7.69 (3.85). Using those figures as the new league average, and refiguring HRAA, RAR, and ARG (RG relative to average), the three come out:

Pujols: 70 HRAA, 79 RAR, 203 ARG
Votto: 63, 72, 194
Gonzalez: 53, 61, 184

I have no choice but to conclude that Pujols was the superior offensive player--to the extent that the tools being used capture reality. You can knock a few runs off of Pujols' figure for excess intentional walks, if you'd like, but it's not enough to make the gap disappear. Factoring in other areas of the game don't figure to do much to boost Votto--Pujols has a good fielding reputation and a track record of good performance in metrics, although this year the two are both rated as just about average by both UZR and RS, with a one run edge for Votto. BP's figures have Pujols as a +5 baserunner, Votto average.

To swing the comparison in Votto's favor, you either need to put stock in a metric like WPA (Votto was +7, Pujols +5.4) or give Votto a bonus because his team bested Pujols' for the division crown. I do neither.

The other interesting comparison is Pujols v. Halladay. Both have 82 RAR initially, but Pujols would actually pick up a few runs for fielding and baserunning, while Halladay would have to lose a tick for his hitting (-1 RC). Factor in the peripheral issue discussed re: Hernandez, and I favor Pujols. This is the second time in three years that I have listed Halladay second on a MVP ballot (last time, in the 2008 AL, he was ahead of the position players but lost out to Cliff Lee).

Adam Wainwright and Ubaldo Jimenez are also deserving of prominent positions on the ballot. Among the down ballot position players, I allow fielding to have just enough influence to push Troy Tulowitzki ahead of Carlos Gonzalez for Most Valuable Rockie, and to put Ryan Zimmerman ahead of some others (Dan Uggla, Jayson Werth, Hanley Ramirez, David Wright, notRyan Howard):

1) 1B Albert Pujols, STL
2) SP Roy Halladay, PHI
3) 1B Joey Votto, CIN
4) SP Adam Wainwright, STL
5) SP Ubaldo Jimenez, COL
6) 1B Adrian Gonzalez, SD
7) LF Matt Holliday, STL
8) 3B Ryan Zimmerman, WAS
9) SS Troy Tulowitzki, COL
10) LF Carlos Gonzalez, COL

Monday, November 08, 2010

IBA Ballot: Cy Young

For the Cy Young award, I generally do not consider hitting, although this is more sheer laziness than any strongly held belief that non-pitching aspects of the game shouldn't count towards the Cy. For the majority of pitchers it doesn't really matter (and fielding is at least included in Run Average, even if jumbled up with the other eight guys' glove work).

My suspicion is that the AL Cy will be the award for which my choices most differ from the sabermetric consensus, as I don't make DIPS metrics a primary consideration. My #1 choice is not one of those differences, though. I'll get into the other candidates a bit below, but assume for the sake of argument that the top two candidates are Felix Hernandez and CC Sabathia. Hernandez bests Sabathia in every single category I list on my pitcher report, albeit not always by significant margins:

* Hernandez pitched more innings (249.2 to 237.2)
* Hernandez had a lower ERA (2.34 to 3.12) and a lower RRA (2.96 to 3.39)
* switching to the more traditional RA estimators, Hernandez had a lower eRA (2.97 to 3.48) and a lower dRA (3.50 to 3.82)
* using batted ball inputs, Hernandez had a lower cRA (3.56 to 3.61) and a lower sRA (3.36 to 3.91)
* Hernandez also had a higher percentage of quality starts, which considering it's quality starts and doesn't include a park adjustment isn't something I'd stress, but he leads Sabathia 85% to 76%.
* So of course Hernandez has the margin in RAA (41 to 28) and RAR (76 to 61)

After Felix, it gets a little less clear--Sabathia leads a pack of five pitchers (Jered Weaver, David Price, Clay Buchholz, and Justin Verlander) separated by just five RAR, with two other pitchers cited as candidates (Cliff Lee and Jon Lester) within another five runs of them. Since a lot of saber-minded people consider similar metrics, I added a dRAR column, based on dRA (my BsR application of DIPS). It requires a new innings pitched figure, dIP, which can be figured as (1 - e%H - %W - %HR)*PA/2.84 (see this post for an explanation of the inputs):



I'm sure you'll see a lot of sabermetric ballots that list Felix #1, but then turn to Lee and Liriano due to their strong showing in the DIPS metrics. For me, they are a secondary consideration, enough to move a pitcher ahead of one a few RAR better, but not enough to turn the ballot upside down. Actual runs allowed contain many biases, but they also carry real and important data (at least from a retrospective value perspective) about sequencing (in addition to the more muffled signals about BABIP). It is also worth noting that when batted ball data is considered (another potential minefield, certainly), a pitcher might give back the advantage dRA indicates (Verlander is the best example here, as his sRA (SIERA-style) is 4.09). Weighing all of the metrics very unscientifically, but giving preference to RAR based on actual runs allowed, this is how I see it:

1) Felix Hernandez, SEA
2) CC Sabathia, NYA
3) Jered Weaver, LAA
4) Cliff Lee, TEX
5) David Price, TB

In the National Leauge, I expect to see much more of a consensus as many of the top candidates have peripherals less impressive than their actual runs allowed rate. Unlike the AL in which pitchers like Lee and Liriano have much better DIPS numbers, in the NL a lot of the top starters move in the same direction. Roy Halladay dRA may be .84 runs higher than his RRA, but Adam Wainwright's is .72, Ubaldo Jimenez's .52, Tim Hudson's a whopping 1.87, Roy Oswalt's .84, Matt Cain's .87, Cole Hamels' .83...this allows us to sideline the ideological debates to a greater extent.

Roy Halladay is the obvious #1 choice, trailing only Josh Johnson in RRA while pitching twenty innings more than anyone else and 67 more innings than Johnson. Not that he should care, but this is actually the first time I've personally ranked Halladay as the top starter in his league. When he won the Cy in 2003, I favored Pedro Martinez or Tim Hudson. In 2005 he was on his way to another Cy Young when he was injured; pitching just 141 innings he still would have ranked second on my ballot. In 2006 he lost a few starts in September and was again second to Johan Santana by my reckoning (although unlike in 2005, Santana was still on pace to earn my vote without Halladay's injury). In 2008 he was second by a slim margin to Cliff Lee, a pitcher he'd eventually become inextricably linked to. In 2009 he had a season so good it would almost always win my vote, but Zack Greinke had to go and turn in the season of the decade.

None of that is to put down Halladay, or say that the 2003 award the BBWAA bestowed upon him was a poor choice (while I favored Pedro, Halladay was a thoroughly defensible choice as well). Rather, it should serve to illustrate how consistently good he's been, and how close he has come to winning three or four Cy Youngs.

After park adjustments, Wainwright and Jimenez are impossibly close--they have the same RA (2.74), nearly the same RRA (2.74 to 2.68), similar ERAs (2.50 to 2.67), the same QS% (76), the same RAA (41) and essentially the RAR (72 to 71). Jimenez looks a little better in the traditional peripheral RAs, while Wainwright looks better in the (not park-adjusted) batted ball RAs. Flip a coin, because you can't go wrong.

Tim Hudson actually has a below-average dRA thanks to a .249 BABIP, but his batted ball metrics look a little better and there's no obvious candidate to replace him. Josh Johnson had an outstanding year, but pitching forty innings less than his competitors consigns him to fifth place:

1) Roy Halladay, PHI
2) Adam Wainwright, STL
3) Ubaldo Jimenez, COL
4) Tim Hudson, ATL
5) Josh Johnson, FLA

Monday, November 01, 2010

IBA Ballot: Rookie of the Year

Over the next few weeks I'll be posting the ballots I submitted to the Internet Baseball Awards, hosted by Baseball Prospectus. While I think too much is made about the post-season awards in general, I also can't deny that they are fun to discuss. Additionally, they present the opportunity to put theories about how to compare player's performance into action, and thus have the potential to stimulate a lot of interesting research and philosophical discussion that can be applied to more general questions (To be clear, that potential has not been fulfilled here.)

I approach the ROY the same way I do the MVP, except limited to rookies of course. I don't consider age, expected future production, or any related factor. I'm perfectly happy to vote for a 34 year old Japanese reliever if they were one of the five top-performing rookies.

Throughout my award posts, the fielding numbers I use are based on the average of Dewan's Runs Saved and Lichtman's UZR, regressed 50% towards zero; or (RS + UZR)/4. Looking at multiple metrics and regressing does not completely alleviate concerns about fielding metrics, but I would not feel comfortable throwing them out completely.

In the AL, the top position player is Austin Jackson, who I have at 27 RAR. Dewan's system loves his fielding (+21); UZR is not as enthusiastic (+4), but it's enough to push him past Brian Matusz for the top spot on my ballot. Matusz had 29 RAR, while Wade Davis had 30, but Matusz' DIPS/batted ball estimators are stronger, and so that puts him ahead on my ballot.

Neftali Feliz is getting some buzz as a candidate, thanks to his saves. I have seen Andrew Bailey's victory last year cited as a reason to support Feliz, but of course, that's a red herring. The comparison should not be of Feliz to a similar past winner, but of Feliz to this year's crop. Without considering leverage, it's not entirely clear that he deserves to rank ahead of another rookie reliever, Daniel Bard. With Jackson at 33 RAR, one would have to give Feliz a leverage multiplier of 1.57 to get them even. His 1.74 LI would suggest a multiplier of about 1.37, which brings him to 29 RAR, roughly equal to Matusz and Davis. I'm still uncomfortable with ascribing that much weight to LI and boosting a reliever who pitched 69.1 IP over a batter with 398 PA playing a demanding position (John Jaso).

Jaso will probably be overlooked by a lot of people, but a catcher with a .376 OBA is nothing to sneeze at. Danny Valencia played well, but he had nearly 70 fewer PA than Jaso, didn't have a significant offensive rate advantage (5.6 to 5.3 RG), and while it doesn't matter retrospectively, his offensive value was largely BA-driven (.313 BA, .205 SEC). I have it:

1) CF Austin Jackson, DET
2) SP Brian Matusz, BAL
3) SP Wade Davis, TB
4) C John Jaso, TB
5) RP Neftali Feliz, TEX

In the National League, the race comes down to Heyward and Posey, so I'll set them aside for a moment to discuss other candidates. Neil Walker checks in at 31 RAR, but -5 fielding and the possible over-adjustment for second baseman in my RAR methodology knocks him off the ballot. Ike Davis was the best rookie first baseman, as far as I can tell, on the basis of his superior OBA to Gaby Sanchez (.358 to .337) and high fielding marks. Chris Johnson's fielding is estimated at -5, which is enough to knock him out of contention, while Starlin Castro's season is more impressive for his age (20) than his performance (albeit not bad at all, 10 RAA and 23 RAR).

Among pitchers, Jaime Garcia stands out at 35 RAR. He will be somewhat overrated by mainstream analysis as just 77% of his runs allowed were earned, the lowest percentage of any NL starter. A 3.60 RRA is quite respectable, though, and his peripherals are similar. Madison Bumgarner was very good as well, turning in 28 RAR in just 110 IP.

I side with Heyward over Posey, largely on the basis of playing time: Heyward played 142 games and batted 611 times, while Posey played 108 games and batted 436 times. It also is important to note that Posey played 35% of his innings at first, which lowers his RAR to 33 versus Heyward's 42. After making that adjustment, Heyward's RG relative to position is 131 versus Posey's 140, which really cuts into Posey's rate stat advantage. Yes, it would have been nice if Posey had spent the whole season in the majors, but Brian Sabean prevented him from contributing for two months, and thus made this an easier choice for me than it seems to be for many others:

1) RF Jason Heyward, ATL
2) C Buster Posey, SF
3) SP Jaime Garcia, STL
4) 1B Ike Davis, NYN
5) SP Madison Bumgarner, SF

Tuesday, October 26, 2010

The Two Best Events in Sports

You can have the Super Bowl, the Stanley Cup, and the NCAA Tournament (except for the games involving my alma mater). Take the Olympics, the Masters, and the World Cup (please take the World Cup, I beg you). Just leave me the World Series and the Breeders' Cup.

Those two events are by far the most compelling (IMO should have gone without saying) in all of sports. Coincidentally, they both occur in autumn, sometimes even overlapping. This year, barring a horrible streak of rainouts, the World Series will have concluded before the horses hit the track at Churchill Downs, but with both events coming up I will bore you with a few of my stray thoughts.

* I do not have a rooting interest when it comes to who wins the World Series, seeing as I don't particularly care for either club. I do have a rooting interest in the series though--rooting for seven games. There has not been a World Series game seven since 2002--also a series that matched San Francisco against an AL West club, apropos of nothing.

If one assumes that the outcomes of each game of a series are independent, and that both teams are equally matched with constant strength and no home field advantage, then the probability of a series of length N is as follows (geometric distribution):

4 = 12.5%
5 = 25%
6 = 31.25%
7 = 31.25%

The probability of going seven years without a seventh game is (1 - .3125)^7 = 7.3%; it's not a particularly likely streak, but it's not remarkable either. I'd still like to see it come to an end in 2010.

* I don't believe there's much value in handicapping a seven-game series, but I would give Texas an edge, something like a 55% chance of winning. There's even less value in doing so on the basis of full-season team records, but I'll proceed for the sake of discussion.

San Francisco had a better actual W% (.568 to .556) and expected W% (based on runs scored and allowed, .581 to .564), but Texas had the edge in predicted W% (based on runs created and runs created allowed, .557 to .543). However, these comparisons don't take into account strength of schedule, which can be a significant factor between the unbalanced schedule and the AL/NL imbalance.

My crude ranking system (yet to be published, as it will take a long, boring post to explain it) gives the Rangers the edge on two of thee comparisons when SOS is taken into account. Based on W/L, Texas has a rating of 121 to San Francisco's 118 (or a 51% chance of winning a seven-game series with no HFA). Based on R/RA, San Fran leads 129 to 119 (54%). Based on RC/RCA, it's Texas 123 to 110 (56%). Considering that, I think 55% is a reasonable estimate.

* From a preseason perspective, when's the last time there was a more surprising World Series than TEX/SF? I intend that as a rhetorical question, as the answer depends on your own perspectives on the teams before the season. For me, it's probably the most surprising since 2005. I picked Texas second in the AL West and San Francisco fourth in the NL West.

There have been other pennant winners that I did not pick for the playoffs (a long list, in fact, owing to both my misjudgments and the inherent inaccuracy of the accuracy), but I had picked both 2005 pennant winners to finish fourth in their division, so that one stands out.

While record in the preceding season is far from a perfect measure of preseason expectations, it might be instructive to look at the combined previous season W% of the two World Series participants. In the expansion era (1961-), the average pennant winner played .550 ball in year X-1. Both TEX (.537) and SF (.543) were below average, although not by a huge margin. Combined, their .540 W% ranks 31st out of the 49 World Series.

Several series in the twenty-first century have been lower, including 2001 NYA/ARI (.533), 2006 DET/STL (.528), 2002 ANA/SF (.509), 2007 BOS/COL (.500), and 2008 TB/PHI (.478). This should not be too surprising, since the expanded playoffs have had the effect of reducing the same season W% of pennant winning teams.

The highest previous season W% of the era was on display in the 1999 NYA/ATL series, by a huge margin; the two teams had combined for a .678 W% in 1998. 1962 NYA/SF (.614), 1970 BAL/CIN (.611), 1978 NYA/LA (.611, and a World Series rematch), and 1964 NYA/STL (.610) are the other high points. The highest of this decade was surprisingly 2003 NYA/FLA; the Marlins were below .500 in 2002, but the Yankees' 103-58 carried the combination to .563.

The two lowest X-1 W% combinations both involve the Twins. Not surprisingly, the worst-to-first MIN/ATL series of 1991 is last at .429, with the 1987 MIN/STL series at .464. The other series featuring teams that had combined to be sub-.500 in the previous season were 1988 OAK/LA (.475), the aforementioned 2008 TB/PHI, 1967 BOS/STL (.478), and 1965 MIN/LA (.491).

The Twins also account for another dubious distinction; their three World Series are the only ones in the expansion era in which both teams were sub-.500 the previous season. Both the 1965 and 1987 series saw Minnesota playing an opponent that had won the NL pennant in year X-2, but struggled in year X-1 before rebounding and taking the flag back.

I'm now descending from "vaguely interesting trivia" to "absolutely worthless drivel", but it's something I noticed looking over the data. This series features one of the closest year X-1 matches between the two participants, with just one game difference between them (TEX was 87-75, SF 88-74 in 2009). The only perfect match of the era is the 1985 STL/KC series (both were 84-78), with 1965 MIN/LA (79-83, 80-82) and 1980 KC/PHI (85-77, 84-78) also off by just one game.

* I have to admit feeling a twinge of happiness with the Phillies' defeat. The Phillies were, both in my estimation and the conventional wisdom, the strongest NL team in 2010. But when a national baseball writer (even if it is a demonstrated fool like Tracy Ringolsby) picks a team to sweep through the playoffs 11-0, it's hard for me to not root against them. It wasn't just Ringolsby--the Vegas notion that the Phillies were 2-1 favorites to win the World Series is tough to defend logically.

The recent Phillies are among the more overhyped teams of recent memory. Their regular season records have been good, certainly, but not historically special. Winning two straight pennants (combined with the third that some members of the media awarded to them) caused a lot of people to downplay the regular season record.

I looked at the team with the best record in the NL over each three-year period beginning in 1961. Obviously, there is nothing special about this approach, no reason to think that looking at three years is better than looking at two or four or using a different approach altogether. It is a timeframe that fits the Phillies' record, as it captures their world title, their pennant, and their best regular season record.

The Phillies' three-year regular season record of 282-204 (.580) is the best in the NL over the last three seasons, but it ranks 28th of 48 in the expansion era, hardly the record of a historically great club. Recent NL leaders with better marks include several combinations of Cardinal seasons (2000-02, 2003-05, 2004-06) and all of the three-year groups formed from Atlanta's 1990s run.

At least to this point, the Phillies' would-be-dynasty is certainly no better than St. Louis' 2004-2006. The Cards record over that period was 288-197 (.594). Their postseason results were the same as the Phillies: a World Series win, a World Series loss, and a NLCS loss.

Absolutely worthless drivel: the best three-year NL record during the period was 310-176 (.638), compiled by the 1997-99 Braves. If you want one which includes a World Series win, it's the 1974-76 Reds (308-178, .634). The lowest three-year mark by a team which led the NL over that stretch is 260-226 (.535) by the 1982-84 Phillies.

* The main storyline for the Breeders' Cup revolves around Zenyatta. For those of you who may be unfamiliar with horse racing, Zenyatta is a six-year old mare that has raced nineteen times in her career and has never be beaten. Nineteen straight is the longest winning streak in major North American horse racing, surpassing the streaks of sixteen compiled by Citation and Cigar. Most of Zenyatta's victories have come in races against other fillies and mares, but after winning the Breeders' Cup Distaff in 2008 (I refuse to refer to this as the "Ladies' Classic" as is now proper), she became the first female ever to win the Breeders' Cup Classic in 2009.

Zenyatta will be retired after the race, and so there would be obvious interest in the final start of a legend, let alone the fact that she could finish her career perfect and become just the second horse to win the Classic twice (Tiznow repeated in 2000-01). She also could become the first horse to win three Breeders' Cup races.

This will be a tough task, however. The Zenyatta-doubters (a group which I admittedly would include myself in) will point out that she has run most of her races over a synthetic surface and in California, and that her only race against males was the 2009 Classic (which, in fairness, is the premiere race in North America).

Zenyatta certainly has a good chance to win, and I can even get behind the idea that she's the deserving favorite. However, if you let me have a choice between Zenyatta and the field, that's easy. Quality Road, Blame, and Lookin' at Lucky all should get support, and there are some other horses of intrigue that may run (like Japanese star Espoir City, the usual European invaders, and second line three-year olds First Dude, Fly Down, and Paddy O'Prado).

* A related Zenyatta storyline that some racing writers have begun to wring their hands over is whether Zenyatta will be Horse of the Year or not. Horse of the Year is voted on by a group of turf writers, similar to the MVP award. However, it's held in a little higher esteem in the horse racing world than the MVP is in baseball circles. A better comparison is college football's national championship, particularly prior to the BCS.

Like the MNC, the winner can be viewed as the overall champion of the season. In college football, you had conference champions (think divisional awards for horse racing, like Champion Three-Year Old Male, Champion Older Female, or Champion Sprinter) and bowl game champions (think Breeder's Cup race winners). There was no unified way to pick an overall champion, so journalists got together and took a poll. The strange thing is that people found themselves intensely emotionally invested in the outcome of that poll, but so be it.

Zenyatta, whose career accomplishments pretty straightforwardly place her among the all-time greats, has never been voted HOTY. Some folks seem to be concerned about this apparent contradiction, a great performer in an individual sport never voted as the best in a given year.

Historically, it is very difficult for a filly or mare to get the nod as HOTY. Since 1971, when the current honors (the Eclipse Awards) were introduced, only four female horses have won the honor: All Along in 1983, Lady's Secret in 1986, Azeri in 2002, and Rachel Alexandra in 2009. Generally, HOTY goes to the top older male horse (which is logical since older male horses are generally the best horses. It's similar to the MNC usually going to the most impressive champion of a major conference). A three-year old male also has a clear path to the award, by scoring impressive victories over older horses (as done by Tiznow in 2000 or Curlin in 2007) or by dominating races against other horses of his generation (Point Given in 2001).

Generally, horses from other groups will only get consideration if there is no clear choice among the top males. Even fillies/mares having undefeated seasons will get passed over in favor of a worthy male (see Personal Ensign in 1988; she even won the Whitney against males but couldn't beat out Breeders' Cup Classic winner Alysheba for HOTY).

That's pretty much what happened to Zenyatta in 2008. She won all seven of her races, but didn't face males or run off of a California synthetic track. Curlin had an impressive season, winning the Dubai World Cup, the Stephen Foster, the Woodward, and the Jockey Club Gold Cup, becoming the all-time earnings leader in the process. He got HOTY.

The 2009 HOTY race was a little less conventional as three year-old filly Rachel Alexandra won the award. Rachel Alexandra had won the prestigious Kentucky Oaks and Mother Goose against other three-year old fillies, but also defeated three-year old males in the Preakness and Haskell and older males in the Woodward. She was not entered in the Breeders' Cup because her owners did not want her to run over a synthetic track; her detractors claimed it was because they wanted to duck Zenyatta.

Zenyatta certainly had a case for HOTY, and would have been a reasonable choice, but I agreed with the decision to side with Rachel Alexandra. Rachel Alexandra won races all over the country; Zenyatta ran only in California. Rachel Alexandra was a dirt horse; Zenyatta stuck to synthetic surfaces, which simply have not yet reached the same level of importance in American racing. Zenyatta's win over males was admittedly in the most impressive race possible, but Rachel Alexandra's three wins over males were all in Grade I races. Perhaps the best point in Zenyatta's favor was that she won at the classic 1 1/4 mile distance, while Rachel Alexandra's longest race was the 1 3/16 mile Preakness.

Looking at the 2010 HOTY race, Zenyatta will win by acclimation if she can repeat in the Classic. She will also be an easy choice if any horse other than Blame, Lookin' at Lucky, or Quality Road win the race. In the event that one of those colts win, though, it would be tough to deny them the award. Each would have a head-to-head win over Zenyatta (and the rest of the field) in the most important race of the year. Blame boasts victories in the Stephen Foster and Whitney; Quality Road in the Donn, Met Mile, and Woodward; and Lookin' at Lucky in the Preakness and the Haskell.

It is possible that Zenyatta could be voted HOTY even in the event that one of her top three challengers wins the Classic. However, I would hope that voters would do this out of a belief that she was the top horse in 2010, not out of a desire to right the historical record. (*) A horse, especially a mare, can easily be considered an all-time great through finishing second three times in the HOTY voting.

(*) If this post wasn't already too long, I would advance a half-baked theory about how this is exactly what happened in 1998, when HOTY voters went for Skip Away instead of Awesome Again, after passing over Skip Away for Favorite Trick in 1997.

Monday, October 18, 2010

Even More Mundane Comments on the Playoff Structure

You certainly don't need me to point out to you that run scoring is down in the playoffs compared to the regular season. I am just going to give you some data on the matter, and half-heartedly explore one possible explanation for why that is.

I figured the RPG (total runs in the game by both teams) for every World Series (through 2008) and a comparison of that to the overall RPG for that season (figured as a simple average of the AL and NL RPG) and the RPG for the two World Series participants (again, a simple average). I have limited the scope to the World Series so that the cross-era comparisons are on more of a level footing.

I've averaged the data by decade (a simple average for 1900-1909, 1910-1919, etc.) so that we can see how it has changed over time:



Frankly, this chart surprised me. I had expected that in recent years the disparity between regular season and World Series RPG levels would have increased, but in fact recent decades are the closest matches for regular season scoring.

Since I assumed this to be the case, I was going to put forth the argument that the playoff structure coupled with changes in the game (specifically, the increased use of relief pitchers) has caused the post-season to become a different game from the regular season, to a greater extent than in the past. My personal take on this phenomenon was to be that it was unfortunate--that the run scoring levels, pitcher usage, strategy choices, etc. should ideally be as close to the same as possible for the regular season and the post-season. I don't like the idea of playing 162 games under one set of conditions and switching to very different conditions to crown a champion.

But my assumption was unfounded. While run scoring declines in the World Series (you can pick your explanation--colder weather, reduced usage of marginal pitchers, increased usage of one-run strategies, or whatever other theory you'd like to advance), the decline has not grown over time. Today's World Series are generally as close to regular season scoring levels as they have ever been.

One other little tidbit to note is that generally, with the 1990s and 2000s actually being the most obvious exceptions, the pennant winners combine for a higher RPG than the majors as a whole. Obviously we expect that pennant winners are very good teams, and will likely both score more runs than the league average and allow less. If a team was equally good offensively and defensively in terms of runs above average, then their RPG would be equal to the league average.

However, if pennant winners were especially strong on defense relative to offense, then their RPG should be lower than the league average. Of course, you can rightly point out that runs scored and allowed have different win values, dependent on their unique combination of runs scored and allowed, and so you don't want to draw too much of a conclusion from this one way or another. Park factors are also ignored by this crude comparison. But if pitching and defense were everything as a minority of traditionalists would have you believe, then pennant winners should certainly have lower RPG than the league average.

Moving along, one possible explanation for lower scoring levels in the post-season is increased usage of top pitchers. I did a little crude investigating on this front by figuring the percentage of regular season innings thrown by a team's top three pitchers (in terms of innings), and comparing that the percentage of World Series innings thrown by the top three pitchers (again, in terms of innings). Please note that I did not consider the same three pitchers--the group under consideration is the three pitchers with the most innings in the games being considered (regular season or World Series).

The reason I chose three pitchers is because presumably the top three in IP will be the front three starters, which is all teams have traditionally needed to use in a seven-game series (of course in today's game four starters are usually employed). There are a number of weaknesses to this approach, including but by no means limited to:

1. It doesn't include the effect of relief aces, who have a disproportionate impact on win probability thanks to working in high leverage situations, and are often employed differently in the playoffs. They are also a relatively modern phenomenon that will damage cross-era comparisons of IP%.

2. It doesn't account for injuries and other factors that alter pitching workloads. If, for instance, a top pitcher is out for the Series, IP% will likely be lower than it might have been, but only because of the absence of the pitcher, not because of any intentional alteration in strategy.

3. IP% can be highly influenced by series length. If a series only goes four games, then it is likely that a larger percentage of the workload can be borne by the key pitchers of the staff.

4. Rainouts or other delays in the series can greatly skew the results by allowing pitchers to pitch more than they would have. This is particularly evident in the 1989 World Series and its earthquake delay; the A's IP% for the series was 88%, the highest in twenty years.

5. I am only considering the World Series; presumably managers are more conservative with the usage of their pitching staffs in earlier playoff rounds, or at least no more aggressive.

So I'm not claiming that these results are particularly informative. Nonetheless, I broke them up by decade as I did for the RPG data. Reg IP% is the simple average of IP% for the two pennant winners, simply averaged for the decade; WS IP% is the same for the World Series; and RAT is the ratio of WS IP% to Reg IP%, expressed as a percentage:



Again, I have to admit this is not what I expected to see. I expected that teams of earlier eras, heavily concentrating their workload on a few pitchers to begin with, would show a more even IP% between the regular season and World Series. The opposite appears to be true; earlier teams ratcheted up the workload for frontline pitchers in the Series to a greater extent than to today's pennant winners. Of course, the weakness in using three pitchers is illustrated by the fact that the ratio was fairly stable until the 1970s, around which time the trend towards larger starting staffs was accelerating.

Again, the data here is by no means conclusive or even particularly insightful. However, I expected to find support for my seat-of-the-pants belief that style of play in the playoffs had become more removed from the regular season over time. Instead, I have no solid ground to stand on to make such a claim (there may well be data out there that would support such a position, but it isn't here).

My argument would have been that since 1) teams could now get a higher percentage of innings from front-line pitchers in the World Series than in the regular season, and 2) that because runs scored declined more precipitously in the World Series, ergo changes should be made to the playoff series format to make it closer to regular season conditions. The most obvious alteration would be to eliminate off-days, which would eliminate the possibility of using a three-man rotation and possibly even encourage the use of five starters as in the regular season.

Leaving aside the practical problems with such a change (chief among them revenue concerns), I personally believe that such modifications would make the playoff series a better test of team strength since they would more closely track the conditions of the regular season. But I didn't find any evidence that the disparity between regular season play and World Series play has increased over time--to the limited extent that the data here addresses the issue, the disparity has actually lessened. Any push for changes that would close the gap is undermined by the fact that larger disparities (at least in terms of these two measures) were accepted throughout the twentieth century.

Tuesday, October 12, 2010

Two Wildcards? Too Many

Apparently, the vast conspiracy that determines what national baseball writers should pontificate about has finally tired of steroids, and has moved on to the pressing issue of whether or not there should be two wildcards in each league. Tom Verducci, Buster Olney, and Jayson Stark have all penned articles on this topic.

I normally ignore the musings of that type of baseball writer, but sometimes it's harder to do that for any number of reasons. It might be that the moral outrage is completely off the scales (as with a typical steroids column) or that the idea is unbelievably stupid (as with the calls for Bud Selig to whitewash events from a game so that Armando Galarraga could be a trivia answer). In this case, not only do I consider the idea stupid, but it would seriously dampen my own enthusiasm for the playoffs.

The folly of wasting one's time on this sort of thing is that just because Jayson Stark advocates something doesn't mean it has a snowball's chance in hell of coming to fruition, and I don't think this proposal is any different. However, in the course of responding I have some potentially interesting data for you on the records of playoff teams in the wildcard era.

In the 32-league seasons since the wildcard was implemented (1995-2010), the average W% for the best team in the league is .620. The second-best division winner averages .583, the third-best .556. The wildcard team is .573 on average, while the team that would be the second wildcard averages .548.

Ten times (31%) the wildcard has had the second-best record in the league, better than every team except the one that bested it for the division title. It has happened 6 times in the AL and 4 in the NL. You might expect that this happens disproportionately when an AL East team wins the wildcard, benefiting the Yankees or Red Sox. That is in fact the case. The wildcard has come out of the AL East twelve times, and in five of those seasons (42%) has had the second-best record in the league. That still leaves five seasons out of 20 (20%) in which the wildcard was not an AL East team and had the league's second-best record.

Only eight times has the wildcard had the worst record of the playoff participants (25%), twice in the AL and six times in the NL. This has become the usual circumstance in the NL, as the wildcard has not bested the #3 division winner since 2004. The opposite holds in the AL, where it has not happened since 1999, when West champ Texas edged out wildcard Boston by one game. In fact, the only other time it happened in the AL was in 1996, when West champ Texas had a better record than wildcard Baltimore.

Of course, these W% comparisons don't account for the strength of the team's schedules, which can be significant in the era of the unbalanced schedule. That would involve some more extensive computations, and I'm not sure it would significantly change the results. I have no doubt that the AL East was easily the strongest division in baseball in 2010, and yet it still managed to produce the wildcard and the team that would have been the second wildcard.

Once one accepts that wildcards are going to be part of the playoff format, I don't think it makes a lot of sense to construct additional barriers to them winning, which is what the second wildcard proposal would do. And while its proponents claim that it would emphasize winning the division, they seem to gloss over the fact that it would inevitably at some point allow a third-place team to qualify for the playoffs? Of course, you could limit the second to wildcard to only second-place teams, but that would further reduce the expected W% of that team and increase dependence on the division format.

And why exactly is winning the division important anyway? Does it really make sense to put more emphasis winning divisions when they consist of uneven numbers of teams and when they are often not even close to being competitively balanced (see the 2009-2010 NL Central)? Why should a future team in the mold of the 2010 Rays have to jump through hoops just because they happen to be a member of the same arbitrary five team grouping as New York? It's bad enough that teams in the West only have to defeat three opponents.

It seems to me as if a lot of the proponents of this plan have really never accepted the expanded playoffs. They yearn for the days in which there were only two divisions, and the possibility existed for two teams with great records to slug it out and one to be shut out. Of course, the reality is that this was less common than some would have you believe (how convenient that the classic ATL/SF race of 1993 happened in the last year of the old format and thus is frozen in time), but they have a point. There is something to be said for using the regular season and not a five-game series to cull the field down to four teams, or even two. I would have no objections with a return to that format.

However, if the expanded playoffs are non-negotiable, then any attempt to punish a wildcard team with an outstanding record only serves to further de-emphasize regular season success in every regard other than defeating the teams in one's own division, while simultaneously giving an opportunity to the other teams that have failed in their divisional races. It does emphasize winning a division, but that division is itself a far cry from the six or seven team groupings that existed in the would-be golden age. And in doing so, it does nothing to emphasize regular season success for teams that are fortunate enough to be grouped with three-five weak teams (hello, 2010 Rangers). The two best teams in baseball slug it out while mediocrities lope comfortably home? Sorry, I don't think that's excitement or upholding tradition--I think it's farcical.