Tuesday, October 28, 2008

IBA Ballot: MVP

Presented below is my ballot (and some justification) for one of the categories in the Internet Baseball Awards hosted at Baseball Prospectus. I’m just one person, and the whole point of having a vote like the IBA is to get a wide variety of (intelligent) perspectives, and so I will not feel in the list bit slighted if you don’t give a flip about this.

In the American League, it was a pretty underwhelming years for position players, at least as far as MVP candidacies go. This results in a large number of candidates but no real standouts. Here is a chart-form look at some of the top candidates. RAA and RAR are against an average hitter at the position, so the column “Def” is estimated of runs saved above average at the position, based on Justin’s stats:

NAME                           RAA                RAR                Def  

Rodriguez 40 59 7

Sizemore 34 58 9

Pedroia 34 58 9

Mauer 38 54 N/A

Hamilton 31 53 0

Roberts 31 51 6

Kinsler 34 50 -6

Markakis 26 50 -4

Youkilis 21 44 4

Granderson 22 43 3

Morneau 15 41 -9

One thing to note is that in choosing my order, I will not treat the fielding stats as 100% equally valuable to hitting; they are pretty clearly less reliable. Also, the approach of comparing to the average hitter can be argued to favor second baseman and shortchange center fielders, which is something to keep in mind. Also, ARod’s “clutch” numbers are dreadful--while I don’t place much weight on this, it’s something that can swing my opinion in the case of a virtual tie.

As a result of all that, I would go with Joe Mauer (who does well fielding, +7 according to Chone’s estimates, although they don’t account for the quality of the pitching staff) in a close race over Sizemore and Pedroia. However, the most generous possible final RAR for Sizemore (giving him the 58 runs for hitting, 9 for fielding, and 5 more to correct for undervaluing center fielders) leaves him at +73; for Mauer, +54 RAR +7 fielding + an indeterminate amount for being a catcher…leave both the pitching duo of Cliff Lee and Roy Halladay.

I usually try to avoid giving my MVP support to a pitcher if there is a very small margin. This is not out of any bias against pitchers being the MVP, but because I am more confident in the sabermetric evaluation of hitters. No fielding or bullpen support to worry about, less nagging questions about “hit luck” and peripherals, etc. But here there is a clear demarcation between the two pitchers and everyone else, and I have to respect that. So this is how I see it:

1) SP Cliff Lee, CLE

2) SP Roy Halladay, TOR

3) C Joe Mauer, MIN

4) CF Grady Sizemore, CLE

5) 2B Dustin Pedroia, BOS

6) 3B Alex Rodriguez, NYA

7) SP Jon Lester, BOS

8) 2B Brian Roberts, BAL

9) CF Josh Hamilton, TEX

10) RF Nick Markakis, BAL

If any of the nimrod crowd at BTF (note that I consider this a subset of the BTF commentariat, not the whole) ever see this, I’m sure they will complain about how there are few guys from the left side of the defensive spectrum and attribute this to VORP and its treatment of DHs. The fact of the matter is, the left side of the defensive spectrum players in the AL just aren’t that good. Here are the leaders in Hitting RAR (not accounting for position) in order by position, without their identities:

5, 8, 9, D, 8, D, 3

I listed seven because I gave three spots on the ballot to pitchers (merit-based; I don’t have a three pitcher quota or anything)--thus there are seven position player spots up for grabs. It should stand to reason that if the best a first baseman can do is seventh, without considering defensive value at all, they are not going to fare very well when you do consider it. The third baseman, the two center fielders, and the right fielder all make my ballot. The two DHs are a guy who played for a bad team (Aubrey Huff, BAL) and a guy who played only 126 games, but was extremely productive when in the lineup (Milton Bradley). In fairness to them, they each played a fair amount in the field, but are considered 100% DH because of the single position adjustment used here (Huff played 24 games at first and 33 at third, Bradley 20 in the outfield corners).

So one could certainly argue that one or both are worthy of a ballot spot--but who are you going to replace? Are you going to argue that Huff was more valuable than his two Oriole teammates who played the field all the time, and also hit well? Are you going to argue that Bradley is more valuable than Hamilton, who could have very easily been reversed in roles with Bradley had Texas felt that would make them a better team? You can, but I’d have a hard time buying it.

In the National League, there is a runaway winner. Albert Pujols led players with 300 or more PA in SLG, RC, RG, and all four of the “above baseline” categories I track. He was second in BA, OBA, and secondary average. He did all this while fielding well (albeit at first base) and helping his team stay in contention all year when many (including myself) thought they’d be bad. And while it was arguably his best season (I think I’d hold out for 2003), he didn’t do anything that was way out of line with his track record.

I realize that you, as an intelligent baseball analyst, realized all that and didn’t need a lecture. But as an intelligent baseball analyst, you probably don’t care much about my MVP preferences in any case.

For the rest of the ballot, Hanley Ramirez’ seemingly improved fielding makes him a clear #2--even if you don’t believe he’s +7 out there as Jin’s numbers do, there would have to be around a dozen run error in that estimate to make me place David Wright or Chipper Jones ahead of him (that’s not to say that there couldn’t be a dozen run error, but I’ll bet against it).

Wright and Jones were #1/#2 on my 2007 ballot, but they will be #4/#3 this year. Wright’s RAR edge over Jones is razor-thin despite having nearly 200 more PA, and the zone data has Jones ahead by ten runs in the field. I find that hard to believe, but I was learning towards Jones anyway. Again, you can’t go wrong with either of them.

Behind them, I slip the top two pitchers in, then go with Berkman on the basis of trusting batting stats, although Beltran and Utley, on the strength of +10 performances in the field, could very well be ahead of them. Jose Reyes rounds out my ballot; Justin has him at -6 in the field, and there are number of guys who you could also make a case for (Giles, Holliday, Ludwick, and McCann among them).

1) 1B Albert Pujols, STL

2) SS Hanley Ramirez, FLA

3) 3B Chipper Jones, ATL

4) 3B David Wright, NYN

5) SP Tim Lincecum, SF

6) SP Johan Santana, NYN

7) 1B Lance Berkman, HOU

8) CF Carlos Beltran, NYN

9) 2B Chase Utley, PHI

10) SS Jose Reyes, NYN

Four Mets in the top ten will rub some folks the wrong way, but I’m hardly the first to observe that New York is a team with several stars and a lot of mediocre filler around them. It is a testament to Wright, Santana, Beltran, and Reyes that they came as close as they did.

Finally, I apologize again for the terrible formatting. Blogger has made it damned near impossible to copy and paste from Word while maintaining a readable output. The "Meanderings" post looked awful and this one may be worse.

Tuesday, October 21, 2008

Meanderings

Here are some disjointed observations and digressions largely inspired by my annual look at the final stats. I have to apologize that they are kind of Indian-centric; I strive to be non-partisan here, but I can’t help that they are the team to which I pay the most attention:

* I want to mention this before the Rays have a chance to ruin it, but if you look at the expansions in groups of two teams, one of the teams has won the World Series and the other has not. This is true for all of the expansions except the 1969 NL expansion, in which neither team has won:

1961: Angels, Senators
1962: Mets, Colt .45s
1969N: Padres, Expos (the exception)
1969A: Royals, Pilots
1977: Blue Jays, Mariners
1993: Marlins, Rockies
1998: Diamondbacks, Devil Rays

Please note that I’m just pointing this out as a coincidence, not any kind of profound insight.

* The AL hit .267/.332/.420, while the NL hit .260/.327/.413. The AL walk/at bat ratio was .096 (.090 with intentional walks removed), while the NL’s was .100 (.091). The AL and NL both had an isolated power of .152. So the biggest real difference in offense between the leagues was seven points of batting average.

Despite this, the AL managed to score .188 runs per (AB - H + CS) while the NL scored just .178. In terms of Base Runs per out, I have the AL at .189 versus the NL’s .183. The apparent difference from the components is not as large as the actual difference. The extra intentional walks could be a factor, but it could be a number of other things and the discrepancy is not particularly noteworthy.

BTW, all of those stats are for the AL and NL offenses. Interleague play makes the issue of league totals a mess as of course there are both offensive and defensive totals, and they no longer are equal on the league level.

* I list three winning percentage categories in my team spreadsheet. The first is regular W%; the second is EW%, which is Pythagenpat; and the third is PW%, which is Pythagenpat based on Base Runs. Teams for which all three figures are close include (these are displayed W%, EW%, PW%) the Cubs (.602, .614, .604), A’s (.466, .470, .470), White Sox (.546, .551, .548), Yankees (.549, .539, .545), and Cardinals (.531, .534, .529). Teams for which there are big differences include the Angels (.617, .544, .519), Braves (.444, .484, .504), and Padres (.389, .416, .453).

Last year there was much discussion about the Diamondbacks, who outplayed their pythagorean expectation to an extreme extent (they won 90 games despite being outscored). This year they had a .506 W% with an EW% of .509.

* The Indians struggled offensively early in the season, and were getting very good starting pitching. Thus the narrative that has been written for the season by the general fan base is that the offense was inadequate (this is not to say that the pitching is being praised; everyone agrees at the very least that the bullpen was dreadful) and the main cause of the team’s .500 season. However, if you look at the season as a whole, the Indians’ were +34 runs versus the league average (park-adjusted) offensively, and +10 defensively. If you look at Runs Created instead of actual runs, then it is +8/+6. The story may have been written in the early part of the season when the Tribe fell out of the race and started selling, but in the end, the run scoring and run prevention were pretty close.

* The Rangers and their opponents easily had the highest scoring level of any team. The RPG in Texas games was 11.53, while the overall MLB average was 9.30. The second-highest was Detroit at 10.36, over a run per game less.

Adjusting for park, the Rangers still lead the way at 11.20, with Detroit still second at 10.36. Toronto ended up with the lowest scoring context either way (8.17 raw, 8.01 adjusted).

* Speaking of Texas, have you noticed how dreadful Luis Mendoza’s season was? I had no idea until I looked at the stats. Mendoza pitched 63 1/3 innings and allowed 61 earned runs for an 8.67 ERA. It’s worse than that, though, as he also was tagged for 13 unearned runs, raising his RA to 10.52. He also inherited 13 runs and allowed 7 to score, so that would be another three runs surrendered.

Park factors help him, a little bit; his adjusted RA is 10.21. His eRA is 7.76, but his dRA is a much more reasonable 4.96. Opponents hit .384 against him when they put the ball in play.

The last pitcher with an ERA greater than 8.00 allowed to pitch more than 60 innings (in fairness, note that Mendoza is just above that cutoff) was Kyle Davies with Atlanta in 2006 (8.38 in the same 63 1/3 IP). The last pitcher to accomplish this with at least half of his appearances as a reliever was Russ Ortiz in 2006 (8.14 in 63 innings, with 26 appearances and 11 starts; Mendoza had 25 appearances, 11 starts). Beyond them, you have Miguel Batista in 2000 (8.54 in 65 1/3) and Benji Sampson in 1999 (8.11 in 71).

All of this added up to -38 RAR for Mendoza, making him the least valuable player in baseball among those who qualified for my spreadsheets. His RAR is overstated a bit by the fact that I lump pitchers into a binary class of starter or reliever with no gray area. Mendoza pitched 45 innings as a starter and 18.3 as a reliever. Thus, weighting the replacement levels by inning, he comes in at -34 RAR, which is still last in the majors by a considerable margin.

* Aquilino Lopez of the Tigers worked in 48 games, all in relief. He inherited 57 runners and allowed 29 of them to score. 1.19 inherited runners/game led all major league relievers, as does (on the trailers list) the -12 runs saved on inherited runners (acknowledging that this is a crude approach that does not consider where the runners are or the number of outs).

* Craig Breslow had a nice season, albeit over just 47 innings, as a lefty reliever for the Indians and Twins. Cleveland claimed him on waivers from Boston near the end of spring training, and pitched just 8 innings before he was let go again. He serves as an illustration of my biggest frustration with Eric Wedge as a manager.

I will tread lightly here, as this criticism is intended more as a fan than an analyst. However, Breslow was allowed to languish in the bullpen for weeks, never entrusted with any high-leverage situation whatsoever. Then, when he did get to pitch, he was not particularly sharp (surprise, surprise). Wedge picks his horses in the bullpen, and then he rides them hard. He doesn’t seem to be able to develop a bullpen in which five or six guys have valuable roles.

In fairness to him, he didn’t have a lot of material to work with this year.

* Most people are aware of the great performance Oakland got out of Brad Ziegler. What I didn’t notice until I looked at the stats was how well Joey Devine pitched for them this year. I would guess I’m not alone in saying that the main thing I remembered about Devine’s short stint in Atlanta was his propensity to allow grand slams. While Devine only pitched 46 innings this year, he was brilliant by any measure (1.41 RA, .60 ERA, 1.16 eRA, 2.39 dRA) and is still only 25. He’s one to keep an eye on for the future.

* About a month ago I wrote about Cliff Lee and his remarkable season in terms of W-L record compared to that of his team. At the time Lee was 21-2 and Cleveland was 71-73. The final tallies were 22-3 for Lee and 81-81 for the team, so his final NW% dipped to .915, still better than Randy Johnson’s .906 in 1995. That Big Unit season was the best that I could find for ten or more wins in my data for Hall of Fame pitchers.

* A hat tip to R.J. Anderson at Beyond the Box Score is warranted here, as he pointed it out a while back, but I thought it to be curious enough to mention again. The perennially disappointing Daniel Cabrera saw his strikeout rate drop to 4.8, which is woeful for a pitcher with his stuff (he didn’t pitch well in 2007 but was still fanning 7.3 per nine innings). I’m not a scout or a PitchF/x-er so I don’t have anything to add beyond that, but maybe there is something not evident in the traditional stats that explains why Cabrera’s career is floundering so.

* Remember when ARod, Jeter, Garciaparra, and Tejada were all AL shortstops? It seems like a long time ago when you look at the sorry crop of 2008. Only three AL shortstops with 300 or more PA were above average hitters: Jhonny Peralta, Derek Jeter, and Mike Aviles. Peralta, though much-maligned by Indians fans, was arguably the AL’s top shortstop in context-neutral terms. That still does not make him a great player (+22 RAA and +41 RAR before taking off up to ten runs for fielding), but you would think that he was below-average and a millstone listening to the talk shows here.

* Your job is to tell me who these players are:

BA OBA  SLG
.294 .346 .403
.268 .355 .415
.263 .327 .346
.260 .334 .427
.263 .317 .361
.282 .316 .439
.243 .326 .357
.240 .325 .303
.237 .339 .359
What is the common thread here? They are all Toronto Blue Jays with 100 or more PA (the stats for those with more than 300 PA are park-adjusted, while the others are not; that’s lazy and sloppy on my part, but irrelevant to the point). For some of the season the Blue Jays had Joe Inglett, John McDonald, David Eckstein, and Marco Scutaro on the roster simultaneously. Any one or two of those guys may be bale to help your team, but what on earth do you need four of them for?

* It’s hard to find a better offensive value match than Jimmy Rollins and JJ Hardy. Rollins had 614 PA, Hardy 621. Rollins made 405 outs, Hardy 409. Each created 91 runs, so Rollins’ RG was 5.71 and Hardy’s was 5.67. Rollins was +29 RAA, Hardy +28. Both were +45 RAR.

* One of these players is considered a MVP candidate, and one was until his team went in the toilet. The other two are well-known, but are often derided for their fielding, which while not great, is not significantly worse than the other two:
PA O RC RG
638 402 101 6.42
670 437 110 6.43
672 428 108 6.43
691 458 111 6.17
They are Pat Burrell, Carlos Delgado, Ryan Howard, and Prince Fielder.

* Here are three AL players:
BA OBA  SLG
.223 .326 .393
.225 .319 .400
.275 .326 .400
Some people still believe that if you have two players with equal OPS, but one has a higher BA, that the one with the higher BA is more valuable. They believe this despite the fact that more sophisticated run estimators show them to be of nearly identical value, with an edge for the lower BA if anything (with the caveat that we are considering a normal environment in the modern major leagues). This is illustrated by these player’s RGs, which are 4.45, 4.51, and 4.43 respectively. Not that I intend this to prove anything, but the players' (R + RBI)/Out are .32, .33, and .31 respectively. (R + RBI - HR)/Out are .29, .28, .27.

You should always remember that if you have identical OPS but varying BA, the player with the lower BA has a better combination of secondary skills. Incidentally, the players are Brandon Boggs, Gary Sheffield, and Billy Butler.

* I have a junkish-stat abbreviated “SU” for Speed Unit. I do not claim it to be better than Speed Score; as a matter of fact, it’s worse. It is based on triples/ball in play, runs/time on base, stolen base percentage, and stolen base attempt frequency. One of the big problems is that I did not cap each component; Curtis Granderson got a 121 last year (it’s supposed to be a 0-100 scale) because he hit a remarkable number of triples. Anyway, take this for what it’s worth. These are the highest and lowest SU by each position in the majors last year:

POS FAST  SLOW
C Rodriguez (56) Varitek/YMolina(25)
1B Berkman(61) Sexson/Aurilia(30)
2B Weeks(79) Kent(30)
3B Figgins(67) Glaus(29)
SS Reyes(92) Eckstein(36)
LF Crawford(85) Cust/Gonzalez(30)
CF Taveras(92) Rowand(32)
RF Span(79) Ordonez/Jenkins(31)
DH Huff(47) Butler(28)
* Finally, the answers to “name the Blue Jay”. In order, they are Joe Inglett, Lyle Overbay, David Eckstein, Scott Rolen, Aaron Hill, Adam Lind, Kevin Mench, Shannon Stewart, and Gregg Zaun.

Tuesday, October 14, 2008

IBA Ballot: Cy Young

Presented below is my ballot (and some justification) for one of the categories in the Internet Baseball Awards hosted at Baseball Prospectus. I’m just one person, and the whole point of having a vote like the IBA is to get a wide variety of (intelligent) perspectives, and so I will not feel in the list bit slighted if you don’t give a flip about this.

In the AL, let’s get it out of the way upfront: it comes down to Cliff Lee and Roy Halladay. We’ll get back to them in a minute.

For the rest of the ballot, there is a pack of starters within ten RAR of each other that I would consider: Jon Lester, John Danks, Daisuke Matsuzaka, and Ervin Santana. Lester leads at +65 RAR, and while his peripherals aren’t as strong as his actual RA, that’s true for the whole group except Santana, who is last at +55 RA, and who is still only about even with Lester in eRA and dRA. Thus, I give Lester the third spot.

Matsuzaka’s odd season has been well-documented; the stats I list don’t capture it, really, although the fact that just 48% of his starts were quality hints at the issues. I give Danks the edge over both Daisuke and Santana.

That leaves the question of relievers; it will come as no surprise to sabermetrically-inclined readers that I am less than impressed with Francisco Rodriguez as a Cy Young candidate. If any reliever deserves that type of recognition, it is Mariano Rivera. He pitched two more innings, with a RA over one run lower, a RRA almost two runs lower, an ERA about .8 runs lower, an eRA almost two runs lower, and a dRA over one run lower. It’s a clearly superior season by any context-neutral measure you’d like to look at. WPA? Rivera leads him +4.47 to +3.33. I would also put Soria, Nathan, and Papelbon ahead of Rodriguez among AL closers. None of that is said to belittle K-Rod; he may not have had a great season, and he may be grossly overpaid in short order, but he’s still quite good, he’s only 26 even though it seems as if he’s been around forever, and he has a fine track record. He’s wasn’t a worthy Cy Young contender in 2008, though.

Lee v. Halladay. Upfront, yes, I am an Indians fan, although I consider myself well below average on the partisan scale. Also upfront, whatever conclusion I draw says nothing to very little about who I feel is a better pitcher--it's solely about who was a more valuable pitcher in 2008. Going backwards or forwards in time, I would take Halladay in a heartbeat.

Halladay’s big advantage up front is 23 more innings; Lee counters with a .42 run edge in RA, and .19 runs in ERA. In terms of eRA, Halladay bests Lee 3.23-3.08, and they are essentially even in dRA (3.33-3.36, advantage Lee). Win-loss record, evaluated superficially against team W%, favors Lee, .915 to .643. I don’t want to go any deeper than that, since it really has no bearing on my choice, but it does tell the story of why Lee will win the actual award.

In terms of value against baseline (based on RA), I have Lee as +80 (replacement)/+51 (average), and Halladay at +77/+44. So my initial inclination is to give the edge to Lee.

One point that has been offered in Halladay’s favor is that his average opposing batter was better. According to Baseball Prospectus’ figures, the average Halladay opponent hit .266/.342/.425 while Lee’s hit .262/.330/.405. Plugging those lines into ERP, the differences are significant--Halladay's opponents had a 5.04 RG versus 4.60 for Lee (for reference, the AL average was 4.78). Of course, as others have noted, this is based simply on the performance level of those batters this year, not their true talent.

However, I think that going too deep into quality of opponent leads to some tricky issues about what constitutes value. You may consider what follows to be a case of paralysis by analysis, but so be it.

If we have two pitchers, one in a tough division (like the AL East) and one in a weak division (like the AL Central), we would expect a random pitcher from the first team to face a tougher average opponent than a random pitcher from the second team. If we assume that the two pitchers’ true talent is the same, and that they each have the same degree of “luck” for lack of a better term, we would expect the second pitcher to allow less runs, win more games, etc. despite the fact that he has pitched exactly as well as the first pitcher.

You can choose to adjust for this--but you can also argue that from a strict perspective of value, those extra wins for the second team are every bit as real. Had the Blue Jays and the Indians been competing against each other for the wildcard, Toronto would not have gotten bonus points in the standings for facing tougher opponents. One can thus argue that Halladay shouldn’t either. For lack of a better term let’s call this the “actual team wins” argument.

Of course, that raises the issue of baseline. I assume that a replacement level starter allows runs at 125% of the league average--but if he faces opponents that are 5% tougher, we would expect such a pitcher to allow something more like 130%. And thus one can also argue that Halladay should be compared to a higher (in terms of RA) replacement level, since what we are trying to measure is the marginal difference between Halladay and a scrub in his circumstances. Inserting a theoretical replacement level into the discussion is a point against the “actual team wins” argument made in the last paragraph. But I don’t consider it a deathblow to that argument.

Anyway, the Indians’ opponents, weighted by games, had a park-adjusted R/G of 4.70 (I did a general park correction, not a specific one based on where the game was played, which would be preferable). The Blue Jays’ have a R/G of 4.79. Weighting by innings, Lee’s opposing teams had a R/G of 4.62, Halladay 4.81. Lee faced opponents at 98.3% the R/G of his team average, Halladay 100.4% of his. If you compare Lee to a baseline .983 times what I was initially using (the league average times 1.25) and Halladay to the same times 1.004, you get each at +78 RAR. (This approach accepts the premise that strength of opposition should only be adjusted for by comparing a pitcher to his team, not to the league).

I don’t know why the results of the opposing slash lines as given by BP and the weighted team R/G given here vary so much. Obviously, one takes into consideration the actual identities of the batters and one does not, but I don’t think that is necessarily a selling point for BP’s approach. What if you studied the opponents and see that more good left-handed hitters got a day off against Lee? If that were the case (I do not know it to be), that would not be something that I would want to hold against his value. And along the same lines, it should be noted that the opposing hitters’ stats don’t account for platoon differences--maybe more mediocre right-handed hitters get an opportunity to play against a tough lefty and drive the pitcher’s composite opponent numbers down.

Again, I’m not saying that either of those scenarios is the case--it would take a lot of digging to figure it out, and I don’t want to get that involved here. But I do not think it is a given that the slash opponents’ figures from BP are the 100% correct choice if one looks to adjust for quality of opponent.

I have already droned on and on about this and I have not even mentioned a key factor like the quality of team fielding. It appears from team DER that this is a slight impairment for Lee vis-à-vis Halladay, but you know how fielding statistics are. The real takeaway from all of this is that there is a tiny margin separating these two. I think either would be an acceptable choice. I am going to take the coward’s way out and choose Lee because that is what the consensus opinion of the masses will be. But if you want to make a case for Halladay, I won’t put up a fight.

1) Cliff Lee, CLE
2) Roy Halladay, TOR
3) Jon Lester, BOS
4) RP Mariano Rivera, NYA
5) John Danks, CHA

The National League race is closer in RAR but in the end, my choice was clearer. Off the bat, no reliever is on my radar; Hong-Chih Kuo led at +27 RAR, followed by Carlos Marmol (+24) and Brad Lidge (+21). Lidge would get the biggest boost from leverage credit, but it’s not enough to get him into contention.

Again, there are two starters that stand out: Tim Lincecum and Johan Santana. For the other spots, Ryan Dempster, Cole Hamels, Dan Haren, and Brandon Webb are my pool of possible choices.

While I would least want Dempster going forward, he did have a fine season, +58/+32, with solid peripherals (3.14 RA coupled with a 3.38 eRA and 3.60 dRA). Cole Hamels is at +27/+56, but he did no better in peripherals (3.45 RA paired with a 3.62 eRA and 4.10 dRA).

Webb and Haren are an interesting pair since they are teammates. Webb pitched ten more innings, but his RA was .18 runs higher, so Haren beats him in RAR +54-+52. Webb was better in eRA (3.20 to 3.51), but dRA favors Haren (3.25 to 3.42). They faced essentially the same quality of opposing batter; .255/.327/.398 for Haren, .254/.325/.393 for Webb. You can basically flip a coin, and mine comes up Webb.

Finally, Lincecum and Santana. Santana pitched seven more innings with a RA .10 higher, an ERA .02 higher, an eRA .64 higher, and a dRA 1.11 higher. Since Lincecum rates even with Santana in the value measures (+72/+43 versus +71/+42) and thrashes him in peripherals, I think he’s the clear choice in the end:

1) Tim Lincecum, SF
2) Johan Santana, NYN
3) Ryan Dempster, CHN
4) Brandon Webb, ARI
5) Cole Hamels, PHI

Thursday, October 09, 2008

Silly Playoff "Thoughts", Vol. 2

DISCLAIMER: This is not an analytical post.

* I was thrilled to see the White Sox dispatched. I like all four of the teams that are left, and will not be bothered by any possible outcome. I’d prefer Phillies/Red Sox, but Dodgers/Rays and Phillies/Rays would be fine too. If it happens to be Dodgers/Red Sox, the Manny stuff will get very stale after about five minutes of FOX pregame coverage. That’s a poor excuse to root against it, though.

* Has it become illegal to have a five-game Division Series? The last to go five was the Angels/Yankees series in 2005. The last LCS to go the distance was the ALCS of a year ago, but that wasn’t a great series for a seven-gamer, as most (Game 2 and the first seven innings of Game 7 as the exception) weren’t nail-biters. And of course 2002 was the last seven game World Series, and 2003 the last to go at least six. As a general fan of the game with absolutely no rooting interest, it would be nice to see a great series. The last two games of the Red Sox/Angels ALDS were a good start, at least.

* I listened to a fair amount of the first round series on ESPN radio, due to other obligations (multiple day games will do that some times). Some general thoughts on the announcers:

Red Sox/Angels: Dan Shulman and Dave Campbell. I think that this was the best of the four crews; Shulman gets a little melodramatic sometimes, but otherwise I think he’s a fine announcer. All color commentators say things I disagree with (as is to be expected), but Campbell comes across as pretty intelligent. He mentioned Baseball Prospectus several times, although unfortunately mostly in reference to the “Secret Sauce”. At least once it was about DER. On the flip side, he is a product of an evil institution (If you don’t know what I’m talking about, don’t worry about it. I’m not going to elaborate).

Dodgers/Cubs: Jon Miller and …I’d have to look it up. Whoever it was obviously did not leave much of an impression for me, but I do like Miller as the play-by-play guy. Of course, I don’t enjoy him on TV because he’s paired with Joe Morgan. The only worse possible team might be Joe Buck and Tim McCarver, but no network would be silly enough to pair those two up…

Phillies/Brewers: Michael Kay and Steve Phillips. Bleh.

Rays/White Sox: Gary Thorne and Chris Singleton. Singleton is one of the first of a generation of commentators for which I am young enough to remember the entirety of their major league careers. Not that I actually remember the glorious details of Singleton’s. Regardless, that’s a plus, I guess; I think that if you are going to have some ex-player to provide insight, then all things being equal it’s better to have a more recent player. However, he used to work for the White Sox and it showed in the broadcast of this series. I don’t mean to imply that he was biased--just that he talked a lot more about Chicago because he’s much more familiar with them. It's up to the listener to decide whether they would rather have a local announcer intensely familiar with one of the teams or a national announcer intensely familiar with neither.

Thorne is much more likeable than Kay, but he’s a horrible radio announcer, because he seems to forget that he’s not on TV. There were multiple times where he omitted pitches (all of a sudden he would say, “the 2-0 pitch is over the outside corner for a strike”--wait, there were two pitches already?) and did not give the location/trajectory of batted balls (“base hit!” or “base hit into right field”--okay, how hard was it hit? Was it down the line, in the hole, or into right-center? Was it hit on the ground, a line drive, a blooper?) On TV, announcers who give you all the details are obnoxious, but the converse is true on the radio.

I can’t imagine it’s easy for guys who go back and forth between the two mediums, and I respect that. It’s still annoying, though. Mike Hegan of the Indians radio network, who comes across as very likeable, is still someone I hate to listen to because he was a TV or TV/Radio announcer for years. He is now solely on the radio, but has yet to tailor his style to fit it.

Tuesday, October 07, 2008

IBA Ballot: Rookie of the Year

Presented below is my ballot (and some justification) for one of the categories in the Internet Baseball Awards hosted at Baseball Prospectus. I’m just one person, and the whole point of having a vote like the IBA is to get a wide variety of (intelligent) perspectives, and so I will not feel in the list bit slighted if you don’t give a flip about this.

In the American League, it should not be much of a debate. Evan Longoria led all AL rookies with +37 RAR. He did this while being limited to 122 games because of injury and early season promotion shenanigans, hitting in the middle of the order for a playoff team, and playing solid defense.

One other position player makes my ballot: Mike Aviles of the Royals. His long-term prospects are not great as he is 27 and hit .325 with a .198 secondary average, but based on this year’s performance, he was +32 RAR.

The rest of my ballot is filled with pitchers, who seem to be often overlooked in ROY discussions (seriously, Alexei Ramirez over these guys?). Brad Ziegler had a great start to his big league career and then started losing steam in September. Regardless, in 59 innings he compiled some eye-popping numbers, like a .99 RRA and a 1.08 ERA. It will be interesting to see how he fares going forward, with his low %H as a warning flag but his extreme groundball, underhand style perhaps marking a pitcher who will be better than his peripherals.

Ultimately, though, second place on my ballot comes down to Armando Galarraga and Joba Chamberlain. Galarraga pitched 178 innings with a 4.18 RA, +36 RAR, but his .246 %H leads to a dRA (the DIPS, BsR-based stat I’m using) of 5.42. Chamberlain pitched just 100 innings, but they were brilliant, with a 2.78 RRA, +28 RAR, and supporting peripherals.

This is a case where the binary pools of starter and reliever that I force pitchers into can be misleading. I use a replacement level of 111% for relievers and 125% for starters, and Joba’s +28 is versus a reliever. However, he pitched 65 innings as a starter and 35 as a reliever (12 appearances as a starter, 30 as a reliever). If I use the weighted average of the corresponding replacement levels (120%), Joba moves up to +33, just three runs behind Galarraga. I think the wide gap in peripherals plus the higher leverage situation Chamberlain faced as a reliever justify bumping ahead. And so this is how I see it:

1) 3B Evan Longoria, TB
2) P Joba Chamberlain, NYA
3) SP Armando Galarraga, DET
4) SS Mike Aviles, KC
5) RP Brad Ziegler, OAK

In the NL, the choice is even clearer. Geovany Soto was a full-time catcher, hitting .280/.359/.494 and, to the extent that you value it, catching for the NL’s top defensive team (the Cubs led at 4.01 RA/G). At +44, he is ten runs in front of the next rookie in RAR, and should be one of the easier award choices in 2008.

Behind him, Joey Votto had a very good rookie season in Cincinnati, creating 92 runs for +34 RAR. The next position player in the RAR ranking was my pre-season choice for NL ROY, Soto’s teammate Kosuke Fukudome. Obviously, he will be getting no votes from the writers, and with his final tally of +15 RAR and -5 RAA, he doesn’t deserve to.

That leaves pitchers to fill out the rest of the ballot. The Tigers may have found a good rookie pitcher in Galarraga, but they also cast one away last winter in the Renteria trade: Jair Jurrjens, a very solid +32 for the Braves. Hiroki Kuroda was +29 and John Lannan +27. Given Lannan’s 5.17 dRA, I don’t see any reason to change the order at all, and so my NL ballot looks like this:

1) C Geovany Soto, CHN
2) 1B Joey Votto, CIN
3) SP Jair Jurrjens, ATL
4) SP Hiroki Kuroda, LA
5) SP John Lannan, WAS

NOTE: This year the IBA limited ROY ballots to three players. I realize that this is how the actual vote is conducted, but I seem to recall that they [IBA] carried it out to five places in the past. Anyway, you got some bonus prattling as a result of my ignorance of this.

Tuesday, September 30, 2008

Silly Playoff "Thoughts"

DISCLAIMER: This is not an analytical post. Nothing in this post should be taken seriously. This post is a waste of your time.

If you are one of those people who likes to bet on football games, and do so based on tips from those radio hucksters, then boy, do I have a World Series tip for you. You know the clowns I’m talking about--the ones that have a super duper lock of the week that you can buy for $10 and talk fast, tossing in a bunch of ridiculous win-loss records (“The Panthers are 3-11 ATS in their last 14 November home games”) as if they are meaningful.

On that level, you should pick the Chicago White Sox to win the World Series this year. Not because of anything they have done on the field, but because they are my least preferred playoff team. The team that I would have least wanted to win of the eight has won it all in 2001, 2002, 2003, and 2005. That’s four out of the last seven years. Not only that, but the Chicago White Sox with Ozzie Guillen as their manager and Nick Swisher relegated to the bench for Ken Griffey and DeWayne Wise are my least favorite playoff team of my time as a baseball fan.

Not only that, but I picked the White Sox to finish fourth in the AL Central this year. My pick of fourth in the AL Central has been a springboard to greatness for the 2005 White Sox and the 2006 Tigers. Not only that, but my fourth place NL Central pick from 2005 won their pennant. It’s getting bad enough that I think I am going to pick the Indians fourth next year just for the heck of it.

My personal preference for this year would go something like this:

1) Brewers
2) Phillies
3) Red Sox
4) Dodgers
5) Rays
6) Cubs
7) Angels
700,000,000,000) White Sox

I really have nothing against the city of Chicago (seriously!), but the Cubs get no sympathy for me because of the fact that they haven’t won since 1908. I just happen to like the other teams better than I like them. I will not be upset if any of those seven teams win.

The doomsday scenario: Cubs and White Sox play for all the marbles as Chicago is awarded the 2016 Olympics and the worse of the two fools wins the presidency. Chicago Uber Alles!

You have just wasted a few minutes of your life if you made it this far; don’t feel too bad, I wasted a few more of mine writing it. Really, though, I am trying to make a point in a roundabout way. I think that a lot of the playoff analysis that you see out there, even from analytical sites, is kind of silly and overwrought.

It’s pretty hard to pick the outcome of five and seven games series contested between two good teams with a great deal of accuracy. That’s not to say that you shouldn’t try to do it, or that it can’t be a fun activity, but I’m not going to join you this time (I have before and reserve the right to do so again). So what you have here is the ultimate anti-analytical approach--who do I want to win? And if you use irrelevant coincidence as the basis for your predictions, the Calcetines Blancas may be your guys.

Monday, September 29, 2008

End of Season Statistics, 2008

Note: This is largely the same explanation as last year; the only significant change is that I switched from FIP to a Base Runs-centric DIPS. Admittedly, this is completely unnecessary, but I decided to use BsR where I could to back up my advocacy for it. It certainly doesn’t hurt, but it is needlessly complicated for that (DIPS) application. I have also added a "R" column which is for rookie. I based this on the list of choices for the IBA Rookie of the Year listed at Baseball Prospectus. I can't guarantee that I marked every rookie (I tried), but I believe I got all the serious ROY hopefuls at the very least.

For the past several years I have been posting Excel spreadsheets with sabermetric stats like RC for regular players on my website. I have not been doing this because I think it is a unique thing that nobody else does--Hardball Times, Baseball Prospectus, and other sites have similar data available. However, since I figure my own stats for myself anyway, I figured I might as well post it on the net.

This year, I am not putting out Excel spreadsheets, but I will have Google Spreadsheets that I will link to from both this blog and my site. What I wanted to do here is a quick run down of the methodology used. These will be added as they are completed; as I post this, there are none, but by the end of the week they should start popping up.

First, I should acknowledge that the primary data source is Doug’s Stats, and that park data for past seasons comes from KJOK’s park database. Baseball-Reference.com and ESPN.com round out the sources.

The general philosophy of these stats is to do what is easiest while not being too imprecise, unless you can do something just a little bit more complex and be more precise. Or at least it used to be. Then I decided to put my money where my mouth was on the matter of Base Runs for pitchers and teams and Pythagenpat. On the other hand, using ERP as the run estimator is not optimal--I could, in lieu of having empirical linear weights for 2007, use Base Runs or another approach to generate custom linear weights. I have decided that does not constitute a worthwhile improvement. Others might disagree, and that’s alright. I’m not claiming that any of these numbers are the state of the art or cannot be improved upon.

First, the team report. I list Park Factor (PF), Winning %, Expected Winning % (EW%), Predicted Winning % (PW%), Wins, Losses, Runs, Runs Allowed, Runs Created (RC), Runs Created Allowed (RCA), Runs/Game (R/G), Runs Allowed/Game (RA/G), Runs Created per Game (RCG), and Runs Created Allowed per Game (RCAG):

EW% is based on runs and runs allowed in Pythagenpat, with the exponent = RPG^.29. PW% is based on runs created and runs created allowed in Pythagenpat.

Runs Created and Runs Created Allowed are both based on a simple Base Runs formula. For the offense, the formula is:
A = H + W - HR - CS
B = (2TB - H - 4HR + .05W + 1.5SB)*.76
C = AB - H
D = HR
For the defense:
A = H + W - HR
B = (2TB - H - 4HR + .05W)*.78
C = AB - H (approximated as IP*2.82, or whatever the league (AB-H)/IP average is)
D = HR
Of course, these are both put together, like all BsR, as A*B/(B + C) + D. The only difference between the formulas is that I include SB and CS for the offense, but don’t want to waste time scrounging up stolen bases allowed for the defense.

R/G, RA/G, RCG, and RCAG are all calculated straightforwardly by dividing by games, then park adjusted by dividing by park factor. Ideally, you use outs as the denominator, but for teams, outs and games are so closely related that I don’t think it’s worth the extra effort.

Next, we have park factors. I have explained the methodology used to figure the PFs before, but the cliff’s notes version is that they are based on five years of data when applicable, include both runs scored and allowed, and they are regressed towards average (PF = 1), with the amount of regression varying based on the number of years of data used. There are factors for both runs and home runs. The initial PF (unshown) is:
iPF = (H*T/(R*(T - 1) + H) + 1)/2
where H = RPG in home games, R = RPG in road games, T = # teams in league (14 for AL and 16 for NL). Then the iPF is converted to the PF by taking 1- (1-iPF)*x, where x = .6 if one year of data is used, .7 for 2, .8 for 3, and .9 for 4+.

It is important to note, since there always seems to be confusion about this, that these park factors already incorporate the fact that the average player plays 50% on the road and 50% at home. That is what the adding one and dividing by 2 in the iPF is all about. So if I list Fenway Park with a 1.02 PF, that means that it actually increases RPG by 4%.

In the calculation of the PFs, I did not get picky and take out “home” games that were actually at neutral sites, like the Astros/Cubs series that was moved to Milwaukee. They simply don’t cause that big of a problem. Suppose Enron Field (I have nothing against corporate stadium names, but I refuse to learn the new ones when they come along) was a perfectly average park in a league in which there are 4.8 runs/game. At 81 home and road games per year, in the previous four years the Astros and their opponents would have scored 3110.4 runs at home and on the road.

If this season, the Astros played four “home” games in an extreme environment in which say 20 runs were scored per game, they would have 819.2 runs added in to the home total. The road games would contribute 777.6 runs to the five-year total. Now, for the five years the Astros’ home games would have a total of 9.70272 RPG versus 9.6 for the road games. The park factor, when fully figured with the regression factor would be 1.0045, when we know that it should be 1.0000. I’m not going to spend too much time worrying about that kind of discrepancy, and that’s a high end example of what the discrepancy would actually be. And I round off to two decimal places anyway, so both would end up 1.00.

Next is the relief pitchers report. I defined a starting pitcher as one with 15 or more starts. All other pitchers are eligible to be included as a reliever. If a pitcher has 40 appearances, then they are included. Additionally, if a pitcher has 50 innings and less than 50% of his appearances are starts, he is also included here (this allows some swingmen type pitchers who wouldn’t meet either the minimum start or appearance standards to get in).

For all of the player reports, ages are based on simply subtracting their year of birth from 2007. I realize that this is not compatible with how ages are usually listed and so “Age 27” doesn’t necessarily correspond to age 27 as I list it, but it makes everything a heckuva lot easier, and I am more interested in comparing the ages of the players to their contemporaries, for which case it makes very little difference.

Anyway, for relievers, the statistical categories are Games, Innings Pitched, Run Average (RA), Relief Run Average (RRA), Earned Run Average (ERA), Estimated Run Average (eRA), DIPS-style estimated Run Average (dRA), Guess-Future (G-F), Inherited Runners per Game (IR/G), Inherited Runs Saved (IRSV), hits per ball in play (%H), Runs Above Average (RAA), and Runs Above Replacement (RAR).

All of the run averages are park adjusted. RA is R*9/IP, and you know ERA. Relief Run Average subtracts IRSV from runs allowed, and thus is (R - IRSV)*9/IP; it was published in By the Numbers by Sky Andrecheck. eRA, dRA, %H, and RAA will be explained in the starters section.

Guess-Future is a JUNK STAT. G-F is A JUNK STAT. I just wanted to make that clear so that no anonymous commentator posts that without any explanation. It is just something that I have used for some time that combines eRA and strikeout rate into a unitless number. As a rule of thumb, anything under 4 is pretty good. I include it not because I think it is meaningful, but because it is a number that I have been looking at for some time and still like to, despite the fact that it is a JUNK STAT. JUNK STATS can be fun as long as you recognize them for what they are. G-F = 4.46 + .095(eRA) - .113(KG), where KG is strikeouts per 9 innings. JUNK STAT JUNK STAT JUNK STAT JUNK STAT JUNK STAT

Inherited Runners per Game is per relief appearance (G - GS); it is an interesting thing to look at, I think, in lieu of actual leverage data. You can see which closers come in with runners on base, and which are used nearly exclusively to start innings. Of course, you can’t infer too much; there are bad relievers who come in with a lot of people on base, not because they are being used in high leverage situations, but because they are long men or what have you. I think it’s mildly interesting, so I include it.

Inherited Runs Saved is the difference between the number of inherited runs the reliever allowed to score, subtracted from the number of inherited runs an average reliever would have allowed to score, given the same number of inherited runners. I do not park adjust this figure. Of course, the way I am doing it is without regard to which base the runners were on, which of course is a very important thing to know. Obviously, with a lot of these reliever measures, if you have access to WPA and LI data and the like, that will probably be more significant.

IRSV = Inherited Runners*League % Stranded - Inherited Runs Scored

Runs Above Replacement is a comparison of the pitcher to a replacement level reliever, which is assumed to be a .450 pitcher, or as I would prefer to say, one who allows runs at 111% of the league average. So the formula is (1.11*N - RRA)*IP/9, where N is league runs/game. Runs Above Average is simply (N - RRA)*IP/9.

On to the starting pitchers. The categories are Innings Pitched, Run Average, ERA, eRA, dRA, KG, G-F, %H, Neutral W% (NW%), Quality Start% (QS%), RAA, and RAR.

The run averages (RA, ERA, eRA, dRA) are all park-adjusted, simply by dividing by park factor.

eRA is figured by plugging the pitcher’s stats into the Base Runs formula above (the one not including SB and CS that is used for estimating team runs allowed), multiplying the estimated runs by nine and dividing by innings.

dRA is a DIPS method (which of course means that Voros McCracken is the true developer), using Base Runs as the run estimator. This is overkill, since a DIPS estimator like FIP will work just fine, but I decided to use Base Runs wherever I could this year. To find, it first estimate PA as IP*x + H + W, where x = Lg(AB-H)/IP. Then, find %K (K/PA), %W (W/PA), %HR (HR/PA), and BIP% = 1- %K - %W - %HR. Next, find estimated %H (which I will just call %H for the sake of this explanation, but it is not the same as the %H displayed in the stats. That is the pitcher’s actual rate, (H-HR)/(estimated PA-W-K-HR)) as BIP%*Lg%H.

Then you use BsR to find the new estimated RA:

A = %H + %W

B = (2*(%H*Lg(TB-4*HR)/(H-HR) + 4*%HR) - %H - 5*%HR + .05*%W)*.78

C = 1 - %H - %W - %HR

D = %HR

dRA = (A*B/(B+C) + D)/C*25.2/PF

Neutral Winning Percentage is the pitcher’s winning percentage adjusted for the quality of his team. It makes the assumption that all teams are perfectly balanced between offense and defense, and then projects what the pitcher’s W% would be on an average team. I do not place a lot of faith in anything based on wins and losses, of course, and particularly not for a one-year sample. In the long run, we would expect pitchers to pitch for fairly balanced teams and for run support for an individual to be the same as for the pitching staff as a whole. For individual seasons, we know that things are not going to even out.

I used to use Run Support to compare a pitcher’s W% to what he would have been expected to earn, but now I have decided that is more trouble than it is worth. RS can be a pain to run down, and I don’t put a lot of stock in the resulting figures anyway. So why bother? NW% = W% - (Mate + .5)/2 +.5, where Mate is (Team Wins - Pitcher Wins)/(Team Decisions - Pitcher Decisions).

Likewise, I include Quality Start Percentage (which of course is just QS/GS) only because my data source (Doug’s Stats) includes them. As for RAA and RAR for starters, RAA = (N - RA)*IP/9, and RAR = (1.25*N - RA)*IP/9.

For hitters with 300 or more PA, I list Plate Appearances (PA), Outs (O), Batting Average (BA), On Base Average (OBA), Slugging Average (SLG), Runs Created (RC), Runs Created per Game (RG), Secondary Average (SEC), Speed Unit (SU), Hitting Runs Above Average (HRAA), Runs Above Average (RAA), Hitting Runs Above Replacement (HRAR), and Runs Above Replacement (RAR).

I do not bother to include hit batters, so take note of that for players who do get plunked a lot. Therefore, PA are simply AB + W. Outs are AB - H + CS. BA and SLG you know, but remember that without HB and SF, OBA is just (H + W)/(AB + W). Secondary Average = (TB - H + W)/AB. I have not included net steals as many people (and Bill James himself) does--it is solely hitting events.

The park adjustment method I’ve used for BA, OBA, SLG, and SEC deserves a little bit of explanation. It is based on the same principle as the “Willie Davis method” introduced by Bill James in the New Historical Baseball Abstract. The idea is to deflate all of the positive offensive events by a constant percentage in order to make the new runs created estimate from those stats equal to the park adjusted runs created we get from the player’s actual stats. I based it on the run estimator (ERP) that I use here instead of RC.

X = ((TB + .8H + W - .3AB)/PF + .3(AB - H))/(TB + W + .5H)

X is unique for each player and is the deflator. Then, hits, walks, and total bases are all multiplied by X in order to park adjust them. Outs (AB - H) are held constant, so the new At Bat estimate is AB - H + H*X, which can be rewritten as AB - (1 - X)*H. Thus, we can write BA, OBA, SLG, and SEC as:

BA = H*X/(AB - (1 - X)*H)
OBA = (H + W)*X/(AB - (1 - X)*H + W*X)
SLG = TB*X/(AB - (1 - X)*H)
SEC = SLG - BA + (OBA - BA)/(1 - OBA)

Next up is Runs Created, which as previously mentioned is actually Paul Johnson’s ERP. Ideally, I would use a custom linear weights formula for the given league, but ERP is just so darn simple and close to the mark that it’s hard to pass up. I still use the term “RC” partially as a homage to Bill James (seriously, I really like and respect him even if I’ve said negative things about RC and Win Shares), and also because it is just a good term. I like the thought put in your head when you hear “creating” a run better than “producing”, “manufacturing”, “generating”, etc. to say nothing of names like “equivalent” or “extrapolated” runs. None of that is said to put down the creators of those methods--there just aren’t a lot of good, unique names available. Anyway, RC = (TB + .8H + W + .7SB - CS - .3AB)*.322.

RC is park adjusted, by dividing by PF, making all of the value stats that follow park adjusted as well. RG, the rate, is RC/O*25.5. I do not believe that outs are the proper denominator for an individual rate stat, but I also do not believe that the distortions caused are that bad. (I still intend to finish my rate stat series and discuss all of the options in excruciating detail, but alas you’ll have to take my word for it now).

Speed Unit is my own take on a “speed skill” estimator ala Speed Score. I AM NOT CLAIMING THAT IT IS BETTER THAN SPEED SCORE. I don’t use Speed Score because I always like to make up my own crap whenever possible (while of course recognizing that others did it first and better), because some of the categories aren’t readily available, and because I don’t want to mess with square roots. Anyway, it considers four categories: runs per time on base, stolen base percentage (using Bill James’ technique of adding 3 to the numerator and 7 to the denominator), stolen base frequency (steal attempts per time on base), and triples per ball in play. These are then converted to a pseudo Z-score in each category, and are on a 0-100 scale. I will not reprint the formula here, but I have written about it before here. I AM NOT CLAIMING THAT IT IS BETTER THAN SPEED SCORE. I AM NOT CLAIMING THAT IT IS AS GOOD AS SPEED SCORE.

There are a whopping four categories that compare to a baseline; two for average, two for replacement. Hitting RAA compares to a league average hitter; it is in the vein of Pete Palmer’s Batting Runs. RAA compares to an average hitter at the player’s primary position. Hitting RAR compares to a “replacement level” hitter; RAR compares to a replacement level hitter at the player’s primary position. The formulas are:

HRAA = (RG - N)*O/25.5
RAA = (RG - N*PADJ)*O/25.5
HRAR = (RG - .73*N)*O/25.5
RAR = (RG - .73*N*PADJ)*O/25.5

PADJ is the position adjustment, and it is based on 1992-2001 data. For catchers it is .89; for 1B/DH, 1.19; for 2B, .93; for 3B, 1.01; for SS, .86; for LF/RF, 1.12; and for CF, 1.02.

How do I deal with players who split time between teams? I assign all of their statistics to the team with which they played more, even if this means it is across leagues. This is obviously the lazy way out; the optimal thing would be to look at the performance with the teams separately, and then sum them up.

You can stop reading now if you just want to know how the numbers were calculated. The rest of this post will be of a rambling nature and will discuss the underpinnings behind the choices I have made on matters like park adjustments, positional adjustments, run to win converters, and replacement levels.

First of all, the term “replacement level” is obnoxious, because everyone brings their preconceptions to the table about what that means, and people end up talking past each other. Unfortunately, that ship has sailed, and the term “replacement level” is not going away. Secondly, I am not really a believer in replacement level. I don’t deny that it is a valid concept, or that comparisons to replacement level can be useful for answering certain questions. I just don’t believe that replacement level is clearly the correct baseline. I also don’t believe that it’s clearly NOT the correct baseline, and since most sabermetricians use it, I go along with the crowd in this case.

The way that reads is probably too wishy-washy; I do think that it is PROBABLY the correct choice. There are few things in sabermetrics that I am 100% sure of, though, and this is certainly not one of them.

I have used distinct replacement levels for batters, starters, and relievers. For batters, it is 73% of the league RG, or since replacement levels are often discussed in these terms, a .350 W%. For starters, I used 125% of the league RA or a .390 W%. For relievers, I used 111% of the league RA or a .450 W%. I am certainly not positive that any of these choices are “correct”. I do think that it is extremely important to use different replacement levels for starters and relievers; Tango Tiger convinced me of this last year (he actually uses .380, .380, .470 as his baselines). Relievers have a natural RA advantage over starters, and thus their replacements will as well.

Now, park adjustments. Since I am concerned about the player’s value last season, the proper type of PF to use is definitely one based on runs. Given that, there are still two paths you can go down. One is to park adjust the player’s statistics; the other is to park adjust the league or replacement statistics when you plug in to a RAA or RAR formula. I go with the first option, because it is more useful to have adjusted RC or adjusted RA, ERA, etc. than to only have the value stats adjusted. However, given a certain assumption about the run to win converter, the two approaches are equivalent.

Speaking of those RPW: David Smyth, in his Base Wins methodology, uses RPW = RPG. If the RPG is 9.4, then there are 9.4 runs per win. It is true that if you study marginal RPW for teams, the relationship is not linear. However, if you back up from the team and consider things in league context, one can make the case that the proper approach is the simple RPW = RPG.

Given that RPW = RPG, the two park factor approaches are equivalent. Suppose that we have a player in an extreme park (PF = 1.15, approximately like Coors Field) who has a 8 RG before adjusting for park while making 350 outs in a 4.5 N league. The first method of park adjustment, the one I use, converts his value into a neutral park, so his RG is now 8/1.15 = 6.957. We can now compare him directly to the league average:

RAA = (6.957 - 4.5)*350/25.5 = +33.72

The second method would be to adjust the league context. If N = 4.5, then the average player in this park will create 4.5*1.15 = 5.175 runs. Now, to figure RAA, we can use the unadjusted RG of 8:

RAA = (8 - 5.175)*350/25.5 = +38.77

These are not the same, as you can obviously see. The reason for this is that they are in two different contexts. The first figure is in a 9 RPG (2*4.5) context; the second figure is in a 10.35 RPG (2*4.5*1.15) context. Runs have different values in different contexts; that is why we have RPW converters. If we convert to WAA, then we have:

WAA = 33.72/9 = +3.75
WAA = 38.77/10.35 = +3.75

Once you convert to wins, the two approaches are equivalent. This is another advantage for the first approach: since after park adjusting, everyone in the league is in the same context, there is no need to convert to wins at all. Sure, you can convert to wins if you want. If you want to compare to performances from other seasons and other leagues, then you need to. But if all you want to do is compare David Wright to Prince Fielder to Hanley Ramirez, there is no need to convert to wins. Personally, I think that stating something as +34 is a lot nicer than stating it as +3.8, if you can get away with it. None of this is to deny that wins are not the ultimate currency, but runs are directly related to wins, and so there is no difference in conclusion from using them if the RPW is the same for all players, which it is for a given league season coupled with park adjusting runs rather than context.

Finally, there is the matter of position adjustments. What I have done is apply an offensive positional adjustment to set a baseline for each player. A second baseman’s RAA will be figured by comparing his RG to 93% of the league average, while a third baseman’s will compare to 101%, etc. Replacement level is set at 73% of the estimated average for each position.

So what I am doing is comparing to a “replacement hitter at position”. As Tango Tiger has pointed out, there is really no such thing as a “replacement hitter” or a “replacement fielder”--there are just replacement players. Every player is chosen because his total value, both hitting and fielding, is sufficient to justify his inclusion on the team. Segmenting it into hitting and fielding replacements is not realistic and causes mass confusion.

That being said, using “replacement hitter at position” does not cause too many distortions. It is not theoretically correct, but it is practically powerful. For one thing, most players, even those at key defensive positions, are chosen first and foremost for their offense. Empirical work by Keith Woolner has shown that the replacement level hitting performance is about the same for every position, relative to the positional average.

The offensive positional adjustment makes the inherent assumption that the average player at each position is equally valuable. I think that this is close to being true, but it is not quite true. The ideal approach would be to use a defensive positional adjustment, since the real difference between a first baseman and a shortstop is their defensive value. When you bat, all runs count the same, whether you create them as a first baseman or as a shortstop.

Figuring what the defensive positional adjustment should be, though, is easier said than done. Therefore, I use the offensive positional adjustment. So if you want to criticize that choice, or criticize the numbers that result, be my guess. But do not claim that I am holding this up as the correct analytical structure. I am holding it up as the most simple and straightforward structure that conforms to reality reasonably well, and because while the numbers may be flawed, they are at least based on an objective formula. If you feel comfortable with some other assumptions, please feel free to ignore mine.

One other note here is that since the offensive PADJ is a proxy for average defensive value by position, ideally it would be applied by tying it to defensive playing time. I have done it by outs, though. For example, shortstops have a PADJ of .86. If we assume that an average full-time player makes 10% of his team’s outs (about 408 for a 162 game season with 25.5 O/G) and the league has a 4.75 N, the average shortstop is getting an adjustment of (1 - .86)*4.75/25.5*408 = +10.6 runs. However, I am distributing it based on player outs. If you have one shortstop who makes 350 outs and another who makes 425 outs, then the first player will be getting 9.1 runs while the second will be getting 11.1 runs, despite the fact that they may both be full-time players.

The reason I have taken this flawed path is because 1) it ties the position adjustment directly into the RAR formula rather then leaving it as something to subtract on the outside and more importantly 2) there’s no straightforward way to do it. The best would probably be to use defensive innings--set the full-time player to X defensive innings, figure how Derek Jeter’s innings compare to X, and adjust his PADJ accordingly. Games in the field or games played are dicey because they can cause distortion for defensive replacements. Plate Appearances avoid the problem that outs have of being highly related to player quality, but they still have the illogic of basing it on offensive playing time. And of course the differences here are going to be fairly small (a few runs). That is not to say that this way is preferable, but it’s not horrible either, at least as far as I can tell.

Given the inherent assumption of the offensive PADJ that all positions are equally valuable, once we have a player’s RAR, we should account for his defensive value by adding on his runs above average relative to a player at his own position. If there is a shortstop out there who is -2 runs defensively versus an average shortstop, he is without a doubt a plus defensive player, and a more valuable defensive player than a first baseman who was +1 run better than an average first baseman. Regardless, since we have implicitly assumed that they are both average defensively for their position when RAR was calculated, the shortstop will see his value docked two runs. This DOES NOT MEAN that the shortstop has been penalized for his defense. The whole process of accounting for positional differences, going from hitting RAR to positional RAR, has benefited him.

It is with some misgivings that I publish “hitting RAR” at all, since I have already stated that there is no such thing as a replacement level hitter. It is useful to provide a low baseline total offensive evaluation that does not include position, though, and it can also be thought of as the theoretical value above replacement in a world in which nobody plays defense at all.

The DH is a special case, and it caused a lot of confusion when my MVP post was linked at BTF last year. Some of that confusion has to do with assuming that any runs above replacement methodology is the same as VORP from the Baseball Prospectus. Obviously there are similarities between my approach and VORP, but there also key differences. One key difference is that I use a better run estimator. Simple, humble old ERP is, in my opinion, a superior estimator to the complex MLV. I agree with almost all of the logic behind MLV--but using James’ Runs Created as the estimator to fuel it is putting lipstick on a pig (this is a much more exciting way of putting it in the 2008 context, don’t you think?).

The big difference, though, as it relates to the DH, is that VORP considers the DH to be a unique position, and I consider DHs as in the same pool as first baseman. The fact of the matter is that first baseman outhit DH. There are any number of potential explanations for this; DHs are often old or injured, hitting as a DH is harder than hitting as a position player, etc. Anyway, the exact procedure for VORP is propriety, but it is apparent that they use some sort of average DH production to set the DH replacement level. This makes the replacement level for a DH lower than the replacement level for a first baseman.

A couple of the aforementioned nimrods took the fact that VORP did this and assumed that my figures did as well. What I do is evaluate 1B and DH against the same replacement RG. This actually helps first baseman, since the DHs drag the average production of the pool down, thus resulting in a lower replacement level than I would get if I considered first baseman on their own. Contrary to what the chief nimrod thought, this is not “treating a 1B as a DH”. It is “treating a 1B as a 1B/DH”.

It is true, however, that this method assumes that a 1B and a DH have equal defensive value. Obviously, a DH has no defensive value. What I advocate to correct this is to treat a DH as a bad defensive first baseman, and thus knock another five or ten runs off of his RAR for a full-time player. I do not incorporate this into the published numbers, but you should keep it in mind. However, there is no need to adjust the figures for first baseman upwards, despite what the nimrods might think. The simple fact of the matter is that first baseman get higher RAR figures by being pooled with the DHs than they would otherwise.

2008 Park Factors

2008 Leagues

2008 Teams

2008 AL Relievers

2008 NL Relievers

2008 AL Starters

2008 NL Starters

2008 AL Hitters

2008 NL Hitters

Tuesday, September 23, 2008

Most Vacuous Post

I hate to even write a post like this because it is the kind of thing that sportswriters love to write about--controversial, timely, and of very little actual importance in the grand scheme of things. So if you’re not interested, don’t read it.

There is a lot of opinion involved with this topic, on all sides. By stating your own opinion, some people will inherently conclude that you are disrespecting others. While there certainly are views on this topic that I don’t particularly respect, because I think they are inane or ridiculous, my intention here is simply to present my position, arguing against others only in the limited way that is required when you try to state your own.

The other problem is that this topic has been discussed by so many people for so many years that there really is nothing new to say. So there’s certainly no claim that any line of argument here is unique.

Nonetheless, here are my thoughts on the criteria for the MVP award, presented here so that I don’t have to cover the topic when I actually pick my IBA ballot.

First, let me reprint here the actual instructions that are sent to each BBWAA MVP voter:

Dear Voter:

There is no clear-cut definition of what Most Valuable means. It is up to the individual voter to decide who was the Most Valuable Player in each league to his team. The MVP need not come from a division winner or other playoff qualifier.

The rules of the voting remain the same as they were written on the first ballot in 1931:

1. Actual value of a player to his team, that is, strength of offense and defense.
2. Number of games played.
3. General character, disposition, loyalty and effort.
4. Former winners are eligible.
5. Members of the committee may vote for more than one member of a team.

You are also urged to give serious consideration to all your selections, from one to 10. A 10th-place vote can influence the outcome of an election. You must fill in all 10 places on your ballot.

Keep in mind that all players are eligible for MVP, and that includes pitchers and designated hitters.

Only regular-season performances are to be taken into consideration.

As Larry Mahnken points out in the THT piece in which the quoted text appeared, the first line basically allows the individual voter to define “most valuable” any which way they want. So how should an individual, whether they are a BBWAA voter or not, define “value”?

That is a question that, to paraphrase someone who is in the news a lot now (and with any good fortune at all will vacate that position in six weeks or so), is above my pay grade. The definition of “value” is by no means clear and is not agreed upon even in the sabermetric community where matters like definitions of terms are taken seriously. In this case, I do believe that it is a matter of personal discretion. Therefore, I will try to explain my personal approach. And it is just that, and if any BTF readers think this is indulgent, then by all means, stop reading now.

The fundamental starting point from which I come from is: The name that was chosen for the award need not be viewed through the lens of how a sabermetrician would define “value” in its most literal sense. This could also be called the “WPA is not the boss of me” principle.

Some people in the sabermetric community seem to see “valuable” and immediately jump to using what I in the past have called “literal value” methods. The most prominent example is WPA, but there are other methods that would qualify.

This response is completely understandable if the name chosen for the award is taken literally. However, I think that it is a little bit silly to read more into the name itself than its creators put thought into it. I seriously doubt that the BBWAA, when deciding to start a year-end award to honor an individual player, put a lot of effort into debating whether it should be called “Most Valuable”, “Most Outstanding”, “Best”, etc. If they did, the instructions that they left for the voters to make their choice leave scant evidence of it, as they are very bland and do not demand any particular viewpoint on what the award represents.

Had the award been called something else, I wonder how the debates about the award down throughout the years would have gone. If it was a “Most Outstanding Player” award, would there be a large group of people claiming that one could only truly be outstanding on a contender? Is the wish to recognize players from contenders a product of the award name, or a product of a natural desire that would manifest itself even if we were discussing the “Top Performing Player”?

In fact, I wonder if even our sabermetric perspective on what constitutes value has not been shaped by the MVP Award and the ensuing debates. I don’t mean to suggest that what is measured by the “value” metrics (WPA, Win Shares, Value-Added Batting Runs, etc. depending on how exactly you define the term) is not meaningful or that they are not worth calculating. However, I don’t dismiss the possibility that people have attached the “value” label to these metrics at least in part due to the fact that they incorporate context in a way that a baseball fan expects thanks to the MVP debates. All I’m suggesting is that the language may have evolved a bit differently in their absence.

Anyway, I choose to not take the “value” literally, as I don’t see compelling evidence that the original intent of the award was to do so, and I don’t see anything in the criteria for the award that compels me to do so. So let me give my take on each one of the criteria:

1. Actual value of a player to his team, that is, strength of offense and defense.

Again, “actual value”, interpreted sabermetrically, can lead to a conclusion at odds with mine. However (and also again), I don’t believe that interpreting the letter of the rules set out by a group of 1930s baseball writers through the eyes of a 2000s sabermetrician is required. For a non-sports analogy that will get me in trouble with some of you (but it’s September of an election year and I don’t do it that often), I believe this is similar to interpreting the Constitution. The original intent of the authors and the commonly understood meaning of the words at the time of their writing trump any modern reading of the document.

Even if I did feel compelled to hold the opposite viewpoint, they had to go mess it up by clarifying the definition of value, and talking about “strength of offense and defense”. What immediately jumps to my mind here is rate statistics, context-neutral or not. I believe that this interpretation is supported by:

2. Number of games played.

I have a rate, and now I have to balance it against playing time, and I feel comfortable in using value above replacement to address the first two considerations, regardless of what you use to fuel the RAR/VORP/WARP/etc.

Personally, I use a linear weights formula based on the player’s overall statistics as the starting point for my RAR estimates. That means that I am not considering the game situation in which his performance occurred. I adjust for park, but only for the value of runs in the park--if there is some other characteristic of the park (like being doubles-friendly or benefiting left-handed pull hitters) that the player is able to exploit, I don’t care. The player is creating actual wins for his team if he his able to do so.

I do tend to consider situational performance as a tiebreaker. If two candidates are very close, and one has a WPA or WPA/LI or VABR clearly superior to the other, then I might bump him ahead. However, I do not use those metrics as a starting point.

Why? Why not? It’s a matter of personal preference. I think that an award that honors the player who demonstrated the most ability (not a precise term for what I described above) is more interesting than an award for the player who had the best combination of performance and timing. And again, I don’t feel that the “valuable” part of the title needs to be interpreted in a sabermetric sense. Maybe you do, and I’m okay with that.

I also do not believe that WPA and similar measures are necessarily the correct way to measure literal value. They are a measure of real-time value. You could also consider a backwards-looking perspective, in which all runs were equally valuable in the end. You could take this viewpoint a step further and claim that any performances in team losses really didn’t have any value, since they were for naught in the end. A forwards-looking definition of value would not measure value in the sense that sabermetricians generally do; rather it would measure what is often referred to as “ability”. The point is that there are a number of different perspectives from which value can be defined, and any number of roads that you can take from that point. I do not accept the premise that WPA is necessarily the best road.

3. General character, disposition, loyalty and effort.

This serves, as Bill James might say, as a BS dump. I’m not a psychologist/sociologist/behaviorist and I try to avoid playing one on the internet, so I’ll give everybody equal marks here in the vast majority of cases. My point is not that you should look only at the statistics--it is just that I will not be bullied into incorporating the common perceptions for individuals on these points (ARod is a choker, Jeter is clutch, Bonds is a cancer, etc.) If you feel you have insight, go ahead and use it. Just don’t feel compelled to follow the crowd, and don’t expect others to treat your opinion as hard evidence.

In defining the two criteria above in terms of numbers, I’m NOT saying that you should take the list of RAR leaders and simply copy it onto your ballot without taking anything else into account. However, I do think it is helpful to use such a list as a starting point, and make adjustments as you deem necessary.

4. Former winners are eligible.
5. Members of the committee may vote for more than one member of a team.

These anachronisms are necessary because at least one earlier incarnation of the MVP award had the opposite rules, and so the voters needed to be reminded of things we all take for granted now.

An ancillary point that has been brought to the forefront this year due to the performance of Sabathia and to a lesser extent Ramirez is how to deal with players who switch leagues during the season. My position has always been that it is an award for NL MVP, and thus only performance that creates value in the National League should count. However, this is admittedly sort of arbitrary, particularly in this brave new world of interleague play. With the lines between the leagues blurred more than ever, and one could argue that some of a player’s performance in the AL (against NL opponents) is providing value by damaging the opponent of his new team. Or one could just argue that holding on firmly to an AL/NL schism is outdated. Nonetheless, I’m sticking with my position, but without a whole lot of conviction and no desire to claim any sort of high ground.

Fianally, a brief bit on the Cy Young and Rookie of the Year awards. It is incredibly hard to find information online about the rules for these awards. I have always approached the Cy Young as being the best pitcher, and not incorporated batting value into the mix. However, I don’t have any compelling reasoning behind this and have absolutely no qualms with the viewpoint of anyone who wants to include non-pitching contributions. You can probably chalk up my exclusion of offense to sheer laziness.

For Rookie of the Year, I have always treated it exactly as I have the MVP, except limited to rookies. I have never made any allowances for age, potential, or players with top-level experience (in modern times, read Japanese veterans, although Negro Leaguers fit the bill in the early years of the award). It seems that the BBWAA has had a little bit of a backlash against Japanese players in the last few years after giving awards to Nomo, Sasaki, and Suzuki, but I see no compelling reason to follow along.