There are a lot of ways to look at team performance based on margin of victory, runs scored in a game, runs allowed in a game, and a like. In this post I look at a few. It is by no means intended to be a comprehensive examination. The data is from Baseball Prospectus, where you can download it very easily and import it to a spreadsheet.
First, let's look at team performance broken down by run differential for the game. I break games into three categories: one-run games, blowout games (margin of five or more), and other games, which for lack of a better team I'll call middle games. You can certainly quibble with the definition of a blowout, but I like the cutoff at five run because it results in the frequency of one-run games and blowouts being approximately equal.
It is necessary to note (and, embarrassingly, I didn't last year) that any run distribution based on games is subject to the bottom of the ninth/extra inning problem--home team scoring is capped in those innings, or home teams don't bat in those innings at all. The analysis that follows assumes that this impacts all teams equally, but that is not necessarily the case. So keep that thought in the back of your mind if you choose to read what follows.
First up are one-run games. There were 656 one-run games in the majors in 2009, out of 2430 total games (27%). Seattle had the highest frequency (34%) and Pittsburgh the lowest (21%). The table below is sorted by the difference between W% in one-run game and W% in all other games:
The eight playoff participants (who I'll use as a stand-in for quality teams in this piece, although it is somewhat circular) had a combined .556 one-run record and a .590 record in other games. These teams played in one-run games 27% of the time, the same as the league average. Six of the eight playoff teams had lower W%s in one-run games than other games (the Angels and Twins were the exceptions, and .016 was the largest difference). It is of course well-known that one-run records are pulled towards .500 for teams of at all observed W% levels.
Again, I've defined blowout games as those determined by five or more runs. There were 686 such games (28%). Kansas City played in the most (36%) and the Mets the least (23%):
Once again, playoff teams played in blowouts with about the same frequency as other clubs (29%). Five of eight had better records in blowouts than other games, with the Yankees (-.016) displaying the largest dropoff. The playoff teams had a composite .619 record in blowouts versus .565 in other games.
You will notice that the Padres had the largest absolute difference by far, going just 10-32 in blowouts but winning a very respectable 54.2% of non-blowouts.
That leaves the "middle" games--and if you have a catchy idea about what to call them, I'd love to hear it--games determined by 2-4 runs. There were 1088 such games, representing 45% of all big league games. Pittsburgh played in the most (51%) and Kansas City the least (38%):
44% of playoff participants' games were middle games, and they recorded a .570 W% compared to .589 in other games. Two of eight (Yankees and Phillies) had better records in middle games, while Colorado was just about even.
Switching gears, here are the frequencies with which teams scored X runs in a game, along with their W% in those games. I'd run a runs allowed frequency list, but of course it will be the same with the exception of complementary W%s:
The mode is three runs (14.2% of games). The "marg" column shows the marginal increase in observed W% for each additional run scored (I cut it off at ten as the frequencies are not very high). You can see that the most valuable run was the fourth, with a jump from a .337 W% to .523. When teams scored three or fewer runs (which occurred 41.9% of the time), they were 393-1644 (.193); when scoring four or more (58.1% frequency), they were 2037-786 (.722).
Last year, I applied a Jamesian concept of Offensive W% based on team run distribution, and will do so again. However, I will not walk through all of the math in detail, and instead will point you to last year's post if you are interested.
The proposed James approach, which I labeled gOW% for "game" Offensive W%, uses the actual W%s for each X runs scored. Of course, one could attempt to model a theoretical W% for each X rather than using the empirical 2009 data. The former would be preferable, but I don't intend this to be taken as a deadly serious exercise so I will stick with the empirical 2009 data, which is subject to the whims of sample size (we are also ignoring the difference in scoring level between the two leagues). I have also lumped all games with ten or more runs scored together at a W% of .955.
Just to make this clear, if a team was shutout 20% of the time, scored one run 30% of the time, and scored two runs 50% of the time (obviously a ridiculous scenario), their gOW% for 2009 would be:
.20(0) + .30(.075) + .50(.208) = .127
I have also figured a standard OW%, holding the defense's RA/G at the league average of 4.61 R/G and using Pythagenpat (there are no park factors applied to any of the results in this post). The traditional OW% and gOW% are usually quite close, as we should expect; when they differ, that means that the run distribution of the team diverged from the expected distribution. Positive differences of gOW% minus OW% indicate teams that bunched their scoring more efficiently than Pyth expected; negative differences indicate an inefficient distribution (keeping the caveat about the ninth inning from the beginning of the post in mind). So, here are a list of teams whose gOW% and OW% differ by more than two games (prorated to 162):
Positive difference: SEA, NYN, SD, CIN
Negative difference: TB
The four teams with higher than expected gOW%s were all bad offenses; none had a gOW% higher than .467. Only one team had a two game negative difference, but the Rays still had a good offense (.521 gOW%), and the next seven teams on the list were all above .500 in that category as well.
We can of course turn this procedure around and use it to calculate gDW%, and standard DW%. The teams with differences of more than two games/162 between the two were:
Positive difference: NYA, KC
Negative difference: LA, PHI, TOR
Finally, we can combine gOW% and gDW% to get what I call gEW%, and compare that to either actual W% or EW% figured from Pythagenpat. The details of the gEW% calculation are given in last year's linked post.
It is not particularly interesting to compare gEW% to actual W%, as most of the biggest differences will occur when teams wins were out of line with expectations based on runs and runs allowed--whether you consider runs in the aggregate (standard EW%) or on the game level (gEW%). Instead, I will list the teams with differences of two games or more per 162 games between gEW% and EW%:
Positive difference: NYA, KC, CIN, HOU, NYN, SEA, SD
Negative difference: LA, PHI, TB, TOR, ATL
Here are the six discussed W% estimates for all teams, sorted by gEW%:
Wednesday, January 27, 2010
Run Distribution and W%, 2009
Monday, January 11, 2010
Hitting by Position, 2009
Offensive performance by position (and the closely related topic of positional adjustments) has always interested me, and so each year I like to examine the most recent season's totals. I believe that offensive positional averages can be an important tool for approximating the defensive value of each position, but they are certainly not a magic bullet and need to include more than one year of data if they are to be utilized in that capacity. So the discussion that follows is not rigorous and focuses on 2009 only.
The first obvious thing to look at is the positional totals for 2009, with the data coming from Baseball-Reference.com. "MLB” is the overall total for MLB, which is not the same as the sum of all the positions here, as pinch-hitters and runners are not included in those. “POS” is the MLB totals minus the pitcher totals, yielding the composite performance by non-pitchers. “PADJ” is the position adjustment, which is the position RG divided by the position (non-pitcher) average. “LPADJ” is the long-term positional adjustment that I used, based on 1992-2001 data. The rows “79” and “3D” are the combined corner outfield and 1B/DH totals, respectively:
I will refrain from commenting, since any drawing any conclusions based on one year of data would be inappropriate.
Next, let's look at NL pitchers' hitting by team. I'm not even going to snark about this, at least not this year. The teams are ranked by RAA above an average pitching staff, which as you can see in the chart above created .28 runs per game. RG is based on just the basic offensive stats plus SB and CS, so sacrifices, which I don't have to tell you are quite frequent for pitchers, play no part in the estimate:
For the second consecutive season, Cubs pitchers were the most productive, but they did drop from +18 to +9. St. Louis, second last year at +8, dropped to +5 but still finished a respectable third. As a group, major league pitchers created runs at just 6% of the overall average, which is about as low as it ever has been.
For the second consecutive season, Toronto pitchers failed to reach base (by hit or walk) in 20 PA, bringing their two year total to 36 PA without reaching. Assuming that they had a true OBA talent of just .150, the probability of this is 1 in 347. Cleveland, Texas, and Minnesota pitchers all failed to collect a hit in their limited chances, but Toronto stood alone in failing to reach base.
Now, let's take a look at the teams with the highest and lowest RAA at each position. RAA is figured using the 2009 positional averages only, without distinguishing between leagues, which means that AL and NL teams will not sum to zero. This is a little distracting but doesn't do much to change the rank order as the opportunities for each team position are relatively equal. Left and right fielders are considered together to figure their baseline (but not first baseman and DHs). The figures have been park-adjusted.
I will not bother running a chart for highest RAA, as it is not exactly surprising that Minnesota catchers or St. Louis first baseman were standouts. It is the trailing teams that are more interesting to investigate. A simple list of the leading teams will suffice:
C--MIN, 1B--STL, 2B--PHI, 3B--TB, SS--FLA, LF--LA, CF--NYN, RF--PHI, DH--NYA
It is interesting that Dodger left fielders and Met center fielders managed to lead the way despite missing their flagship players for significant portions of the season. The Dodgers had Manny (944 OPS in 418 PA) and Juan Pierre (778 in 308) in left almost exclusively (only 4 PA went to other hitters), and it is obvious that their first place showing was mostly Manny's doing. Carlos Beltran was the main force behind the Met performance (930 OPS in 342 PA), but Angel Pagan (818 in 264) helped the cause as well, and 109 PA came from four other players.
Now, the worst performances, along with the player who led the team in games played at the position:
AL teams do not fare well, but it's more of a coincidence than a systematic bias--AL position players (not including pinch-hitters) combined to hit .268/.333/.430 while their NL counterparts hit .267/.336/.425, essentially the same level of production. It just so happened that Kansas City and Seattle punted on getting offensive out of two positions each.
Let me conclude by looking at the RAA for each position, with negative performances in red and those +/- 20 runs bolded (an arbitrary cutoff, to be sure). The teams are grouped by division and sorted by the sum of RAA across all listed positions:
The Phillies led the NL with six above average positions. Strangely, first base, manned by perennial MVP candidate Ryan Howard, was only their fourth-most productive spot. At least the Mets can't blame their star players for team shortcomings this year, although I'm sure someone will try. As you already knew, Washington's offense wasn't really that bad, but in this division that leaves them at the bottom.
Milwaukee picked up the Mets' banner for a star-driven offense this year, leading the NL with three +20 RAA positions, although second base has to be considered a surprise. They led the NL in infield and corner (1B, 3B, LF, and RF) RAA. On the other hand, the Cards were operating on the one star plan. The Pirates had just one above average position, tied for the fewest in the NL, and were last in the NL in infield and corner RAA, but still beat out the Reds, the only NL team with four -20 positions and the owners of the worst outfield production in the circuit. It's a good thing they got rid of Adam Dunn (*) and signed Willy Taveras, obviously.
(*) I'm not making any statement on Dunn's fielding or his contract, just the amazing ability of some Reds fans to scapegoat him for everything, including his offense.
Los Angeles boasted baseball's most productive outfield, but was just +2 total at the other positions. For the second straight season, the Padres got surprising production out of center field; last year it was Jody Gerut and Scott Hairston who manned the position. This year, Hairston was great in 146 PA and 378 PA from Tony Gwynn Jr. were good enough. Arizona and San Francisco both have offenses filled with black holes, but the DBacks had five above average performers while only the Giants third baseman (read: Sandoval) were above average.
Another entry in the obvious department: the Yankee offense was good. They led the AL in above average positions (8), infield RAA, and corner RAA. All four infield positions were +20 or better, and their center fielders just missed breaking even. Boston had the AL's top outfield RAA, with solid performances everywhere except shortstop.
Minnesota did their best to balance out Joe Mauer's amazing season with a disastrous second base performance, while Cleveland joined the Angels and Pirates as the only teams in baseball with no positions +/- 20 runs. The Pirates were bad overall, the Indians average, and the Angels good. Detroit and Kansas City were the only teams in the AL with just two positive positions, but Chicago wasn't far behind as none of their four above average positions were better than +4. The Royals had four -20 positions and the worst outfield RAA in the majors.
I didn't make any attempt to measure it formally, but the Angles probably had the most balanced positional offense in baseball, with all positions in the -6 to +19 range. Ranger first baseman had the lowest RAA of any position. Oakland had the majors' lowest infield and corner RAA, largely driven by their first baseman, who were almost as bad as those of the Rangers. Seattle was worse overall, despite a good performance from Ichiro in right, as they had three black holes.
Yuniesky Betancourt deserves a special mention, as he helped lead the Royals to the lowest RAA at short in the majors and the Mariners to second worst. He didn't do it all alone, however; shockingly, in both stints he had a higher OPS at short than his teammates did. Taking 39% of the Mariners SS PA, he had a 609 OPS vs. 587 for his teammates, and in 43% of the Royals SS PA he OPSed 642 to his teammates 517. And that's just sad.
Here is a link to a spreadsheet with each team's performance by position.
Monday, January 04, 2010
Hitting by Lineup Slot, 2009
This piece has next to no analysis--it is mostly a presentation of data that you could easily get elsewhere. But since I devote an entire post each year to the most productive leadoff slots in the majors, I've decided to also make one post dealing with the other eight lineup positions. You wouldn't want the #7 hitters to feel neglected, now would you?
First, let's take a look at the average production out of each lineup slot, broken down by league. All data has been culled from Baseball-Reference. An important technical note upfront: I have chosen to use the standard ERP formula to estimate runs created for all lineup positions. In reality, of course, the appropriate linear weight values vary by context, one of which is most certainly batting order position. Folks like Tango Tiger have done a great deal of research on LW by batting order, and while the differences in event coefficients are not monumental, it would be more accurate and more interesting to use them. For this post, though, the weights are the same for each position:
NL #3 hitters blow every other position out of the water, boasting stars like Pujols, Utley, Ramirez, Braun, Gonzalez, etc. amongst their ranks. Not surprising is the fact that NL #9 hitters are the least productive, and that amongst spots manned (nearly, thanks to Tony LaRussa) exclusively by real hitters, the #8 batters from both leaguers bring up the rear.
Making conclusions from one year of data of this type is dangerous, but I found it interesting that #2 hitters in both leagues out-produced the league average. Historically, #2 hitters have often been below average in OPS+, although I should point out that the figures here do include runs produced through stolen base attempts. The top producing AL lineup slot was the cleanup hitters, although the #3 and #5 spots were essentially as productive.
Let's next take a look at the top ranking team (in RG) for each lineup slot. The player listed is the most common batter in that spot, by games started:
First, I should acknowledge that the most common batter can be misleading, as in the case of Tampa's #6 hitters. Pat Burrell did appear in the most games in that spot (47), but his OPS was just 764. The team was actually propelled to a strong showing by the performances of Ben Zobrist (26 G, 1147 OPS), Willy Aybar (28, 999) and Gabe Gross (15, 919), as well as productive eight game stints by Evan Longoria, Carlos Pena, and Jason Bartlett.
The Yankees were often cited as having a "deep" lineup, and being the most productive at three spots backs up those claims.
And the trailers:
How bad were the Kansas City cleanup hitters? So bad that only two slots (excluding NL #9s) produced lower RGs: Seattle #7 and Detroit #9. They were last in BA, second last in OBA, and only Seattle #7, San Diego #8, and their teammates who batted ninth compiled a lower SLG.
The Seattle #7s were an embarrassment in their own right--they were a point behind Colorado's #9 hitters in OBA (although park adjustments would mercifully rescue them if applied).
I also figured runs above average versus the league average for the lineup slot (AL and NL separate, no park adjustments). These figures are available in the spreadsheet at the end of the post. I need to emphasize that they compare to the league average for 2009 only, and so they are subject to the players actually batting in particular spots. That's a clumsy way of saying don't take them too seriously. AL #3 and #4 hitters each created 5.5 runs per game, but the NL breakdown was 6.4/5.9. If a given NL team's third hitters and cleanup hitters created 6.2 runs, then the #3s would be ranked below average and the #4s above average. But does the fact that Pujols, Ramirez, Utley, and others bat third rather than fourth really change the value of another team's #3 hitters? No.
Anyway, caveats aside, here were the ten highest RAA figures from individual slots:
Don't worry: David Wright was more responsible for the Mets' #5 hitters than Jeff Francoeur, but Francoeur appeared in more games.
The bottom ten:
Boston and Colorado had the most above-average lineup slots, with eight (Red Sox leadoff and Rockies #7 were the below-average performers). Detroit (#4 the exception), Oakland (#9), Pittsburgh (#1), and San Diego (#9) all had eight below-average slots.
There are a lot of other ways you could look at this data; I'll leave you to it if you want, as I've run out of interesting things to say:
http://spreadsheets.google.com/pub?key=tUplx96l-xDtgKWCJmdXXqw&output=html
Monday, December 21, 2009
A Caution on the Use of Baselined Metrics per PA
I threw this together based on a Twitter discussion I had last week that included Justin (@jinazreds), Matt (@devilfingers), Josuha (@JDSussman), and Erik (@Erik_Manning)--hopefully I didn't miss anybody. Said discussion was going just fine between the other parties until I stepped in and said the opposite of what I meant, so I need to clarify my point. The end result will be that I obfuscate my point, but that's par for the course around here.
I really should just get around to writing the rate stat series that I have been promising since I started this blog, and then I could give my thoughts on this topic from A-Z in one place. But this is a lot easier, and the rate stat series would have eight parts and be remarkably dry reading as I go around in circles.
Suppose we want to express a baselined measure of value as a rate stat. In this case, I'll work with something similar to Palmer's Batting Wins--wins above average, considering only offensive production--but the theory behind it has wider applications.
The standard way of doing this (incidentally, one of the few things that Tango Tiger, David Smyth, and myself ever fully agreed upon on the topic of rate stats in our many discussions at FanHome (at least at the time--I certainly don't presume to speak on behalf of those gentleman)) is to look at BW/PA. If we were working with a standard runs created method, we would look at RC/out. But when our metric has already been baselined to average, we have already incorporated the run value of avoiding outs/generating PA. RAA/Out will double-count that aspect of offense, more or less.
Of course, we all recognize that the value of a run varies depending on the context in which the hitter plays, so we convert RAA to WAA, and we have something like Batting Wins. Let's look at two players credited with a similar number of BW, but in very different contexts with a big difference in PA:
Nap Lajoie, 1903 AL: 5.8 BW in 509 PA
Frank Thomas, 1996 AL: 6.2 BW in 636 PA
Incidentally, the BW figures here are my rough estimates; for the purposes of this discussion, it doesn't really matter how they reflect specifically on Lajoie and Thomas--I don't care to compare them to see who was better, I just needed a good example. They actually differ fairly substantially from those published elsewhere, but that's not important. There will also be some rounding discrepancies from using just one decimal place throughout the post, but the purpose of this exercise is not a precise examination of the two players.
Figuring BW/650 PA, we come up with Lajoie at 7.4 and Thomas at 6.3. From this, we can conclude that Lajoie was significantly more productive on a rate basis as an offensive player, right?
Let's get a second opinion first. If the stat we wanted to put on a rate basis was standard Runs Created, we'd generally do that by taking RC/Out and comparing it to the league average. My estimates have Lajoie at 207 and Thomas at 195. One needs not be an expert on the relationship between the scales of the two metrics to realize that 207-195 is a much narrower gap than 7.4-6.3.
What is the cause of this discrepancy? It's not the RC/RAA inputs, since they are based on the same formulas. It's not a case of the metrics being incompatible--RC/Out and RAA/PA (or BW/PA) correlate very highly when the samples are drawn from similar contexts.
The problem is that Plate Appearances (which are obviously the denominator for BW/PA) are not constant across contexts. Outs are, more or less. No matter what era the game is played in, what park it's played in, how many runs are scored, or anything else, there are still three outs per inning. And (approximately) 27 outs per game. Even if you had five inning games in one league and thirteen inning games in another, it will all wash out (or close to it) when you look at runs per out.
On the other hand, plate appearances are not constant across environments. In 1903, AL teams averaged 35.8 PA/G (actually AB+W only), while in 1996 AL teams averaged 38.7. Therefore, 650 PA in 1903 are not equivalent to 650 PA in 1996. 650 PA in 1903 represent the number than an average offense would generate in 18.2 games, but in 1996 they represent just 16.8 games worth.
Getting back to the actual PA used by Larry and the Big Hurt, one would think that since Thomas came to the plate 127 more times that he had participated in a much larger share of his team's PA (even when we recognize the difference in schedule length). However, Lajoie's 509 PA are equivalent to 14.2 games; Thomas' 636 to 16.4 (*). Thomas had 15% more opportunities when you adjust for context, versus 25% more when only raw PA is considered (and this is without considering the difference in season length).
In a higher PA environment, players will get more raw opportunities, but each PA has less impact on wins and losses, as each represents a smaller portion of a game. We can adjust for this by normalizing Plate Appearances to some "reference level", common for all leagues.
So let's instead look at BW/650 PA, except we'll normalize PA to an average of 37.2/game (this is roughly the post-1901 major league average). Lajoie will now be credited with (5.8/509)*(35.8/37.2)*650 = 7.1 BW/650 and Thomas with (6.2/636)*(38.7/37.2)*650 = 6.6 BW/650. The gap is .5 BW, whereas before normalizing PA it was 1.1.
If you'd like a formula:
(Baselined metric/PA)*[(Reference PA/G)/(League PA/G)] = baselined metric/normalized PA
or
baselined metric/normalized PA = [(Baselined metric)*(Reference PA/G)]/[PA*(League PA/G)]
where "reference PA/G" is simply the fixed PA/G value everything is being scaled to (37.2 in the Lajoie/Thomas example)
When looking at players within the same league, one doesn't have to worry about this issue--in that situation, one doesn't even have to convert from runs to wins unless they are so inclined.
Let me circle back and explain the underlying premise of this post again, as I'm pretty sure I've been too verbose and may have distracted from it. Basically, the point I am trying to make is that a batter's contribution occurs within the context of his team's games (or, if we'd like to divorce the player from his actual team, the idealized games of a league average team). What matters is not the raw number of plate appearances a batter gets, but the proportion of his team's plate appearances that he gets. That's the point, in a nutshell.
So we could look at Lajoie/Thomas from that perspective as well, making it explicit with the use of percentages. Lajoie played in a league in which there were 140 games in a season and 35.8 PA/G, so the average team would get 140*35.8 = 5,012 PA, of which he was given 10.2% (509/5012). Thomas was given 10.1% of the idealized team's PA (636/162/38.7).
Therefore, their opportunity as measured in PA was essentially equal. Thomas actually had 127 more plate appearances because he played in an environment in which there were a lot more to go around in each game, and because he played in a league in which there were 22 extra games played. We want to adjust for the former cause when looking at BW/PA; the latter is not a problem because Thomas also had 22 extra games in which to increase his raw number of BW (it might be something you want to consider, in Lajoie's favor, if you are comparing raw BW totals).
(Incidentally, one can use this principle to try to adjust for the differing numbers of PA players get as a result of being on good or bad offensive teams, even within the same league. The most notable metric to incorporate this factor is David Tate's Marginal Lineup Value. I'll leave a full discussion of the pros and cons of that approach for another time).
When expressing individual batter's productivity as a rate, there are legitimate reasons not to use outs. I've written about some of them before. The good news, though, is that using outs does not cause an excessive amount of distortion on the player level, as long as you don't take it too far (as Bill James' old system of Offensive Won-Loss Records did). If I had to present just one rate stat and it had to be the most accurate estimate of individual offensive performance I could possibly offer, it would not be runs/out--it would be something like the WAA/Normalized PA presented here or something even more complex. (Just to be clear: if you use outs as a denominator, the numerator should be absolute runs; if you use PA as a denominator, then you can put your baselined metric in the numerator).
But the nice thing about working with outs (and I am fully aware that I'm repeating myself) is that outs are constant across all contexts. Outs are fixed at three per inning whether you play in the Baker Bowl in 1930 or in Dodger Stadium in 1968. Avoiding a lot of headaches that come from making sure you've considered all of the variables when using PA as your denominator might well be worth the tiny bit of distortion that comes with using outs. I know it is for me.
(*) If you really want to get cute, you could argue that we want to look at PA/Out as the number of outs is not constant across all league-seasons due to factors like extra inning games, home teams that don't bat in the bottom of the ninth, rainouts, etc. I wouldn't waste my time but I wanted to acknowledge it.
Tuesday, December 15, 2009
Leadoff Hitters, 2009
Once again, here is a look at the composite performances of the players who batted in the leadoff spot for each team. The data is from baseball-reference.com and again, it includes ALL of the PA out of the leadoff spot. In parentheses I list the players who appeared in twenty or more games in the #1 slot (which is not the same as starting twenty games; they could have been pinch runners, defensive replacements, etc.), but that does not in any way mean that they are the only contributor to the team total.
I always feel obligated to point out that as a sabermetrician, I think that the importance of the batting order is often overstated, and that the best leadoff hitters would generally be the best cleanup hitters, the best #9 hitters, etc. However, since the leadoff spot gets a lot of attention, and teams pay particular attention to the spot, it is instructive to look at how each team fared there.
The conventional wisdom is that the primary job of the leadoff hitter is to get on base and score runs. So let's start by looking at runs scored per 25.5 outs (AB - H + CS):
1. NYA (Jeter), 6.6
2. LAA (Figgins), 6.4
3. TOR (Scutaro), 6.3
7. LA (Fucal/Pierre), 5.9
Leadoff average, 5.3
ML average, 4.6
28. CIN (Taveras/Stubbs/Dickerson), 4.6
29. NYN (Pagan/Reyes/Cora), 4.5
30. OAK (Kennedy/Cabrera/Sweeney), 4.4
I will always list the top and bottom three, as well as the leader and trailer in each league if they are not already included. There will be some different names popping up on the leader lists, as there were a number of changes involving top leadoff hitters: injury-riddled seasons for Jose Reyes and Grady Sizemore, the flip-flop of Johnny Damon and Derek Jeter, and Hanley Ramirez' move into the #3 slot in the Florida batting order.
Next up is the other obvious metric, On Base Average, which here excludes HB and SF:
1. NYA (Jeter), .398
2. LAA (Figgins), .389
3. SEA (Suzuki), .382
6. PIT (McCutchen/Morgan), .362
Leadoff average, .344
ML average, .330
26. OAK (Kennedy/Cabrera/Sweeney), .320
28. SF (Velez/Rowand/Winn), .304
29. CIN (Taveras/Stubbs/Dickerson), .301
30. PHI (Rollins), .293
Two things jarred me when looking at this list--first, the fact that Pirates leadoff hitters led the NL in OBA. Andrew McCutchen (.366 in 487 PA) and Nyjer Morgan (.351 in 211) both contributed to this feat. Meanwhile, on the other side of the state, Jimmy Rollins led the Phillies to baseball's worst mark.
What I call Runners On Base Average is a modified OBA, equal to the Base Runs A factor per PA (or regular OBA less HR and CS in the numerator). It measures the number of times a player is actually on base available to be driven in by a teammate. It penalizes homers, obviously, but if you believe that the role of a leadoff hitter is to get on base for others, that is not necessarily a drawback. The leaders were:
1. NYA (Jeter), .364
2. LAA (Figgins), .359
3. SEA (Suzuki), .355
4. STL (Schumaker/Ryan/Lugo), .348
Leadoff average, .313
ML average, .296
28. CIN (Taveras/Stubbs/Dickerson), .272
29. DET (Granderson), .266
30. PHI (Rollins), .256
The Tigers leadoff men led baseball with 34 homers, dropping their already below-average .321 OBA to last in the AL when homers are removed. Incidentally, Astros leadoff hitters hit the fewest longballs (4).
Runs to RBI ratio is not a measure of quality, but rather of shape. The conventional stereotype of an ideal leadoff man would have a high ratio; those who are non-traditional are more likely to have a low ratio:
1. CIN (Taveras/Stubbs/Dickerson), 2.5
2. STL (Schumaker/Ryan/Lugo), 2.4
3. WAS (Guzman/Morgan/Harris), 2.4
5. LAA (Figgins), 2.1
Leadoff average, 1.6
ML average, 1.0
28. TEX (Kinsler/Borbon), 1.3
29. DET (Granderson), 1.2
30. SF (Velez/Rowand/Winn), 1.2
As you can see with just a glance, R/RBI ratio does not track the quality measures above very closely. Cincinnati ranked in the bottom three in the first group of metrics we examined, but here they lead the way, not due to any particular ability to score runs but due to their anemic .348 SLG (last) and .093 ISO (third last, ahead of only HOU and LAA). The Angels rank high as well, yet did well in runs scored and OBA.
Bill James' designed his Run Element Ratio for a similar purpose--identifying whether hitters fit the traditional mold of table setters or cleanup men. RER is the ratio of steals and walks (both events that do little to advance other baserunners) to extra bases (power). We should expect somewhat similar results to R/RBI ratio, but without the influence of teammates and with singles excluded from consideration:
1. LAA (Figgins), 2.4
2. HOU (Bourn/Matsui), 2.0
3. BOS (Ellsbury/Pedroia), 1.4
Leadoff average, 1.1
ML average, .8
28. PHI (Rollins), .7
29. SF (Velez/Rowand/Winn), .7
30. DET (Granderson), .6
Another Bill James measure was what I'll call Leadoff Efficiency--an estimated runs scored per 25.5 outs. James' formula assumes that 35% of runners on first (estimated as S + W - SB - CS) will score; 55% of runners on second (D + SB); 80% of runners on third (T); and of course homers always result in a run scored. As Tango Tiger has pointed out here in the past, these weights are not particularly accurate, which is evidenced by the fact that the average LE is 6% higher than the average of actual runs scored/25.5 outs for leadoff men. Nevertheless, it is James' metric and I'll present it as he figures it:
1. NYA (Jeter), 7.3
2. SEA (Suzuki), 6.4
3. TOR (Scutaro), 6.3
5. PIT (McCutchen/Morgan), 6.3
Leadoff average, 5.7
ML average, 5.5
28. OAK (Kennedy/Cabrera/Sweeney), 5.0
29. SD (Gwynn/Cabrera), 4.9
30. CIN (Taveras/Stubbs/Dickerson), 4.6
Transitioning back to metrics that are designed for more general application, David Smyth has suggested using 2*OBA + SLG for leadoff hitters. Since the most accurate weight for OBA in an OPS-type construction (for the purpose of predicting team runs scored) is somewhere in the vicinity of 1.5-1.8, using a weight of two gives a little bit of a boost to OBA, but not excessively so (and still closer to the ideal weight than what is used in standard OPS or even OPS+). I have taken 70% of the result to bring it back onto the normal OPS scale; since neither OPS nor 2OPS is on an organic scale, we might as well stick with the more familiar scale:
1. NYA (Jeter), 892
2. SEA (Suzuki), 851
3. TOR (Scutaro), 816
5. PIT (McCutchen/Morgan), 811
Leadoff average, 769
ML average, 754
27. OAK (Kennedy/Cabrera/Sweeney), 705
28. PHI (Rollins), 701
29. SD (Gwynn/Cabrera), 694
30. CIN (Taveras/Stubbs/Dickerson), 665
Finally, we can always just evaluate a leadoff hitter in the same way we'd generally evaluate any other: standard Runs Created per Game:
1. NYA (Jeter), 7.1
2. SEA (Suzuki), 6.2
3. PIT (McCutchen/Morgan), 5.7
Leadoff average, 5.0
ML average, 4.8
28. OAK (Kennedy/Cabrera/Sweeney), 4.1
29. SD (Gwynn/Cabrera), 3.8
30. CIN (Taveras/Stubbs/Dickerson), 3.7
If writing a piece like this obligates one to anoint one team's leadoff men as the most effective, then it's the Yankees, led by Derek Jeter. The worst? Well, it's tough to believe, but Willy Taveras managed to do what Jerry Hairston, Corey Patterson, and friends could not in 2008--lead the Reds leadoff slot to the bottom of the rankings in three categories.
Here is a link to a spreadsheet with all of the data, sorted by OBA:
Leadoff Hitters 2009
Wednesday, December 09, 2009
(Informally) Grading BBWAA Award Choices
Last time I tried to explain why I don't particularly care about whom the BBWAA annual awards are bestowed upon, and how my feelings on those awards differ from those I hold on some other awards.
Now I'm going to turn around and talk about the very results I claimed not to care about, which will understandably lead to charges of wanting to have it both ways. Perhaps, but I hope that my previous missive will allow you to see where I'm coming from.
First, a brief digression. While my opinion is of course the one I value most, I am nowhere near vain enough to assume that you care about my opinion (while also recognizing that I am not infallible). So I do take a look at the Internet Baseball Awards, now maintained by Baseball Prospectus, and add my two cents into that voting. I believe that we yahoos on the internet, as a group, do make better choices than the writers do as a group. Are there IBA results that I personally find dubious? Of course, but I think that overall they are more sensible than what the BBWAA proffers.
For fun, I am going to propose a series of letter grades by which to judge the BBWAA awards against your own judgment. I will illustrate this by looking at the MVP winners for the last ten seasons, and comparing my choices to those of the BBWAA and the IBA. I have also limited myself in making my selections only to what I felt at the time. I have not gone back and reviewed the statistics (or some of the new data that has become available, like better fielding metrics) to see if I would still view those awards the same way I do now. Remember, I'm not claiming that my opinion is infallible, and I certainly wouldn't make my claim about what my opinion was ten years ago. Also, the frequency of the grades doesn't make an abundance of sense--A+ is more common than A, for instance. The point here is just to offer a systematic way of categorizing your *own* opinion on the outcome of the vote, with mine just serving as a superfluous example.
The first letter grade is A+ (I've avoided pluses and minuses except in this case, as they are needlessly complex for a silly application, but you could figure out how to mix them in elsewhere if you wanted). An A+ selection is one that you agree with--the singular choice of the BBWAA is the singular choice that you would have made. The last BBWAA A+ selection (in MVP voting and in my opinion, of course, which will go unstated for the rest of the piece) were Albert Pujols and Joe Mauer in 2009.
An A selection is one in which you would have chosen a different player, but could have yourself made the case for the actual winner. Your candidate and the winner were very close and while you went one way, you wouldn't even waste your time trying to dissuade someone that endorsed the other player. The last A selection for me was Albert Pujols, 2005 NL. I felt that Derrek Lee was a sliver more valuable, but it was hard to argue that with any conviction.
A B selection is one win which you have a clear preference for a different candidate, but you can certainly see why others might support the winning player. This player will probably be in the top five on your ballot (or top three for the Cy Young), and his value estimate should be close enough to that of your player that it is within a reasonably restrictive confidence interval. The last B selection was Dustin Pedroia, 2008 AL. While I felt that one of the top two pitchers (Lee or Halladay) should have won the award, and that Mauer or Sizemore were more deserving position players, Pedroia was hardly an outlandish pick. I didn't endorse him, but it was a solid selection.
A C selection is one where you feel the player was clearly inferior to another, and while he would have been on the bottom of your ballot (or just off of it in the case of the Cy Young), you have a hard time accepting him as the best choice. The last C selection was Jimmy Rollins, 2007. I had Rollins eighth on my ballot, and felt that David Wright and Chipper Jones stood out as the top two. I also had Rollins behind two other players at his position and one other player on his team; he had a fine season, but the MVP was a bit much.
A D selection occurs when you don't feel the player should have even been in the top ten. This will likely only happen when the mainstream evaluation of the player's statistics differs widely from the sabermetric evaluation, or when the media has latched onto a storyline about a particular player and built an MVP case around it. In the last ten years, there has not been a D selection, only because of the (possibly too) large definition I have assigned to grade F.
A F selection is the same as a D selection, except the player is also judged to be inferior to one or ideally two or more comparable players. I used three criteria for comparable:
1) a teammate
2) a player at the same position and a somewhat similar profile as a hitter (Mark Grace would not be comparable to Frank Thomas, even though they were both first baseman; Jim Thome would be)
3) if the winner came from a contender, then a comparable player under condition #2 must have also come from a contender
The last F selection was Justin Morneau, 2006 AL. I believe that Morneau was not one of the ten most valuable players in the American League AND that his case was inferior to that of his teammate Joe Mauer.
I hope I've made it clear that I don't intend this exercise to be taken too seriously; it is just an organized way of assessing how the actual award choice compares to your own. It turns out that, even under the light of the grading system, the MVP choices have been decent for the last ten years. It's been even better in the NL, largely due to the presence of two superstars that are hard to ignore (although the AL does have an answer in Alex Rodriguez).
However, for my money the results of the IBA balloting have been nearly flawless. Only twice in twenty votes did I feel that there was a demonstrably more deserving recipient--and in both of those cases, I accept that it is possible that the IBA winner was truly the MVP under my personal standards (grade B choices). Sixteen times I have agreed with the IBA choice (A+), while three times it has been too close to call and I went with the other good option (A).
The uncharitable way of looking at this would be to say that I am a stathead ideologue, and that the other IBA voters (since they are self-selected among folks who at least have exposure to sabermetriclly-aware outlets) are ideologues as well, and so it is no surprise that there is a consensus. Perhaps. I tend to think that it illustrates that an informed, diverse group can make excellent decisions and arrive at consensus through the power of logic and analysis. But in the end it's all just for fun, so that would be a bit far to push it.
Tuesday, December 01, 2009
The MVP, the Hall of Fame, and the Emmys
In the past I have written disdainfully of the BBWAA post-season awards, going so far as to say that I don't care. I've said the same thing about Hall of Fame voting.
Whenever I do this, the post seems to get linked somewhere and people ask "If you don't care about it, why are you writing about it?" It's true that "I don't care" is a fairly strong declaration, and that what I'm actually aiming for is "I don't care about the specific outcomes of the voting process. I am interested in ways in which the outcomes could be improved by changing the process or the voter pool". Of course, if you need to slap a title on your blog post, the former is a lot easier to work with than the latter. In any event, if you're not interested in my opinion, that's fine by me. Don't read it.
To belabor this point, let me give you an example by discussing four sets of awards/honors that I don't care about in one way or another: the Daytime Emmys, the Primetime Emmys, the BBWAA awards, and the Baseball Hall of Fame. The exact manner in which I don't care about each differs, and should be illustrative of what I'm getting at:
The Daytime Emmys--I don't care about the Daytime Emmys because I don't watch daytime television. Not only does the identity of the award winners have no impact on me, I know and care next to nothing ("next to" is a necessary qualifier to avoid a gotcha when it turns out I've heard of some soap operas) about what is being honored. I don't know who won the awards, I don't care to know, and I don't have any opinion about who should have won them.
The Primetime Emmys--I may not care who wins the awards, but I watch some of the shows eligible for consideration or know something about the others that I don't watch. I'm not a TV critic and make no claims to be one; I watch what I enjoy, and I don't care whether it is considered worthy of praise by critics or considered to be garbage. While I think that it would be cool if LOST won the Emmy for best drama every year (or Monk and/or The Office for best comedy), I can't say that Mad Men is unworthy, because I don't watch it, know little about it, and I don't evaluate TV shows in the same way that Emmy voters do.
Baseball Hall of Fame--Last year I wrote a couple of posts titled "Why I Don't Care About the HOF". The main point was that I don't care about specific Hall of Fame selections (i.e. "Should Blyleven or Trammell be in?" or the endless Jim Rice debates) because I believe the system is too far gone. There have been so many mistakes made that even a concerted effort going forward will not salvage the Hall of Fame as a means to honor truly great players. Additionally, I believe that one of the reasons for the mistakes is the haphazard means of selecting players that have been employed over the years, and the lack of a coherent vision for the player selection process when the institution was founded.
The concept of a Hall of Fame in general, and how a hypothetical one should be constructed, is of interest to me. And so I do offer comments from time to time on how I feel the current Hall could be improved (although this hypothetical improvement would still be insufficient to salvage the inductee roster at this point), or about how a Hall could be designed in theory.
BBWAA Awards--I think that the questions posed by each of these awards are interesting, and I follow the game closely enough to come to my own informed judgments about which player should win. I think the voting process (ten-man ballot, two voters per city in the case of the MVP) itself is solid. I'm not wild about the instructions laid out for voting, but they could certainly be worse. Most importantly, I think it's worthwhile to honor the best players of each season
However, while the voting process and instructions are okay, I don't hold the judgment of those doing the voting in particularly high esteem--particularly with respect to a number of de facto criteria have emerged (or seem to have emerged). Most prominent amongst the de factor prerequisites I find objectionable are that a player must play for a contender (or otherwise have a clearly superior season to anyone else) and that starting pitchers are not seriously considered. With respect to Rookie of the Year voting, sometimes writers apparently can't be bothered to ascertain which players actually are rookies. And there is the issue that people who will report the news are called on to make the news, which may not have a tangible impact on the voting but raises a red flag just a little bit up the pole.
So at the end of the day I have enough qualms about the BBWAA awards to be uninterested in the results of who wins, except to the extent that the results give us insight into how the voters view the game or how the selection process could be improved. If I feel player X is undeserving, yet he wins the award, I might chuckle and shake my head; I might accuse the voters of overlooking one facet of the game and overvaluing another; but I'm not outraged. I'm not going to write about how Player Y who I prefer was robbed of the award; instead, I'll write about why Player Y really was the most valuable player of the league, which is a question that may be raised and brought to the forefront by the BBWAA awards, but could easily exist in a vacuum (if you think this distinction is splitting hairs, I disagree but understand where you're coming from).
Comparing the Hall of Fame votes to the annual award votes, I prefer the latter. The voting process is designed better, but more importantly, the mistakes of the past only cast a small shadow on present results.
Silly choices by the BBWAA for MVP or Cy Young can set a precedent, to a limited extent. One could attempt to justify voting for a closer as MVP because Willie Hernandez won, or for a player solely on the basis of impressive home run and RBI numbers because of Andre Dawson, 1987. And poor choices, even those in the past, can serve to reduce the respect given to the award.
However, in the case of the Hall of Fame, the mistakes of the past are never far from discussion, since each election builds on the one that came before it. The awards slate is wiped clean each year, but each Hall candidate is compared not only to their ballot mates but to the previous inductees. No single voter is compelled to change his standards to fit previous choices, but comparison to past inductees is unavoidable. And while the impact of a single questionable selection can be minimized (Jim Bottomley doesn't come up much in Hall discussions), a series of questionable selections is harder to push aside (like the Frankie Frisch-era VC selections that Bottomley was a part of). Furthermore, the honor of being a Hall of Famer itself is cheapened by poor selections, as the honor is to be considered in a group with the past inductees.
To summarize, in order to flesh out what I mean when I say I don't care about a certain baseball award, I've offered four gradations of indifference:
1. I care about neither the mission of the award nor the entities being honored (Daytime Emmys)
2. I care about the entities to some extent, but not about the mission of the award (Primetime Emmys)
3. I care about the entities, and think the mission of the award is solid in theory, but the implementation is such that it has lost me other than as a theoretical exercise (Baseball Hall of Fame)
4. I care about the entities, and the mission of the award, but the people entrusted with bestowing the award severely dampen my enthusiasm (BBWAA post-season awards)
Tuesday, November 17, 2009
IBA Ballot: MVP
Disclaimer: Presented below is my ballot (and some justification) for one of the categories in the Internet Baseball Awards hosted at Baseball Prospectus. I’m just one person, and the whole point of having a vote like the IBA is to get a wide variety of (intelligent) perspectives, and so I will not feel in the list bit slighted if you don’t give a flip about this. You've been warned. Also, the RAA and RAR figures that will be cited are my own estimates, detailed here. Any Leverage Index, WPA, or UZR figures cited are from FanGraphs; any quality of opposition or baserunning figures are from Baseball Prospectus.
The AL MVP debate will not be much of a debate after all--with the Twins' September surge, Joe Mauer should coast to the award. As you will see, I ultimately agree with this, but I think there's a solid case to be made that Zack Greinke was the most valuable player in the American League. The statistical comparison between the two hits on any number of hot spots--pitcher v. hitter, DIPS and fielding support, evaluating fielding, what the most appropriate baseline is--and depending on the judgment calls you make on those matters, it is not that hard to come down on Greinke's side.
RAR favors Greinke, +91 to +82. Mauer is generally considered a solid defensive catcher--let's call it five runs in lieu of a more rigorous estimate. On the other hand, that RAR figure assumes that Mauer is a full-time catcher, when in fact he appeared in 109 games behind the plate and 28 as a DH. That knocks around three runs off his position adjustment, leaving him at +84 (please note that I am overstating the precision of the initial estimates and the subsequent adjustments for the sake of discussion). BPro estimates his non-SB baserunning at -3 runs, which would lower his RAR accordingly.
Greinke's RAR is based on just taking his actual runs allowed into consideration. Suppose that you were to use his dRA (basically, simple DIPS RA) as the fuel for RAR instead. In that case, he would drop to...you guessed it, +84. Greinke allowed a high BABIP (not really a surprise with KC's poor fielding behind him), but DIPS throws the situational pitching baby out with the fielding bathwater.
There's also the matter of baseline. If you use average, Mauer is ahead +67 to +61 before considering his defense. If you use something in the middle, you're liable to end up with another statistical tie.
I'm not going to try to argue for one or the other, just that they're too close to call. The deciding factor for me is that Mauer is a position player and Greinke is a pitcher. I have no problem voting for a pitcher for MVP--my ballots probably average around 2.5 pitchers per league season. But if a pitcher and a position player are in a dead heat, I'm going to side with the position player more often than not. Last year I went with Cliff Lee for AL MVP as no position player turned in a comparable season.
Behind them, Roy Halladay and Felix Hernandez had seasons that would often be good enough to win Cy Youngs, and the rest of the AL hitters collectivley had another year without any real jawdropping performances. So the two hurlers go 3-4, with Ben Zobrist and Derek Jeter the next two position players.
Why Zobrist over Jeter? Zobrist does well in the defensive metrics, but you don't have to put a lot of weight on that to make a reasonable case for him over Jeter. I have Zobrist and Jeter even as offensive players without considering position (60 to 59 RAR, Zobrist's superior rate balanced by Jeter's extra 100 PA). So you only have to believe that Zobrist's fielding was more valuable than Jeter's, not that it was truly spectacular.
Evan Longoria's +52 RAR leave him down the ballot if you go just by hitting, but of course he has a good defensive reputation and his UZR was a whopping +19. Even if you only want to credit him as a +10 fielder, it's enough to vault him past some not particularly impressive fielders.
After yet another pitcher (Verlander), the last two spots on the ballot go to first baseman--Mark Teixeira and Miguel Cabrera. Kevin Youkilis might be the most surprising omission from my ballot, and you can certainly make a case for him over either of those two. Even giving him credit for his time at third base, I have him at +48 RAR versus +55 for Teixeira and +53 for Cabrera. My RAR figures lazily omit hit batters, but giving him another three runs for getting plunked and two runs for fielding (Fangraphs' estimate) leaves him in a dead heat. I went with the other two, but reasonable people will surely differ on this one.
Kendry Morales, on the other hand, will get mainstream MVP support but at +42 RAR, he's well behind the other first baseman, and even a generous (and likely unwarranted) fielding estimate just gets him into the mix. Was he a better value than the man he replaced? Absolutely. But I can't call him a more valuable player.
Victor Martinez ranks fourth in RAR among position players, but doesn't crack the ballot. Why? For one thing, the aforementioned RAR figure treats him as a pure catcher, but in reality 46% of his games played were at first base or DH. Incorporating that into his positional adjustment drops his RAR to +48, thirteenth in the league.
1) C Joe Mauer, MIN
2) SP Zack Greinke, KC
3) SP Roy Halladay, TOR
4) SP Felix Hernandez, SEA
5) 2B Ben Zobrist, TB
6) SS Derek Jeter, NYA
7) 3B Evan Longoria, TB
8) SP Justin Verlander, DET
9) 1B Mark Teixeira, NYA
10) 1B Miguel Cabrera, DET
In the National League, there is one super candidate with no real competition. Despite tailing off a bit in the second half, Albert Pujols recorded what is IMO the best season of his career (although picking between Pujols seasons is like picking between...nah, I'm bad at analogies), finishing second in BA, first in OBA, SLG, secondary average, Runs Created, and all four of the baselined categories I track. His RAR lead is a whopping 21 runs over Hanley Ramirez, and there's no amount of finessing the numbers that will close that gap.
Behind him, it is too close to call between Hanley Ramirez and Chase Utley once you give Utley credit for fielding and getting hit...Ramirez is +80 RAR, but you can't give him a big fielding number, while Utley is +64 with a very believable +12 UZR and some runs lying around from plunkings and baserunning. I went with Ramirez because I trust the offensive numbers more, but I wouldn't argue one bit if you think Utley was more valuable. Utley's oft-overlooked contributions allowed him to pass the two big first base bats, Prince Fielder and Adrian Gonzalez, but they are next on my ballot, with Gonzalez getting a narrow edge due to his fielding prowess (he trails 77-74 in RAR).
Ryan Zimmerman had a +18 UZR, which at full credit would put him ahead of the first baseman. I hedge a little bit and place him behind them, followed by a cavalcade of pitchers and Troy Tulowitzki:
1) 1B Albert Pujols, STL
2) SS Hanley Ramirez, FLA
3) 2B Chase Utley, PHI
4) 1B Adrian Gonzalez, SD
5) 1B Prince Fielder, MIL
6) 3B Ryan Zimmerman, WAS
7) SP Tim Lincecum, SF
8) SP Chris Carpenter, STL
9) SP Adam Wainwright, STL
10) SS Troy Tulowitzki, COL