I will admit up front that I have not paid much attention this year to the award debates, either in the mainstream or the sabersphere. This is good in the sense that I am coming into this cold, without having read many other perspectives that might bias me one way or another. It’s also bad for the same reason--while I don’t think I’ve ever found mainstream commentary on player value particularly useful, there are a lot of others out there worth reading.
I simply decided this year that I wasn’t going to waste any time thinking about awards until the season was over. Not that I ever obsessed over them previously, but I pretty much completely shut them out of my mind this year. So much so that when an acquaintance who knows I’m a baseball nut asked me who I thought should be the NL Cy Young winner a couple weeks ago, he was shocked when the best I could offer was “uh, probably either Halladay or Kershaw”.
In any event, let me start in the AL. This rookie crop belongs to the pitchers. My top three candidates are all starting pitchers. Michael Pineda got out the best start, Ivan Nova had the flashiest win-loss record, but Jeremy Hellickson was the AL’s most valuable rookie pitcher. Hellickson led the trio in innings (189 to Pineda’s 171 and Nova’s 165) and RRA (3.24 to 3.81 and 4.10). Combining the two, I have Hellickson at 52 RAR, Pineda 36, and Nova 30.
Hellickson’s BABIP was just .229, so from a strict DIPS perspective one could make a case for Pineda (or even Nova) ahead of Hellickson. But for a retrospective award, I stick to actual runs allowed and first-order component RA for the most part. If Pineda and Hellickson were close, I would consider moving the former ahead, but the gap is too big in this case.
For the remaining two spots on the ballot, the top position players are Dustin Ackley, Eric Hosmer, and Jemile Weeks. Ackley was the most productive hitter of the three, while Hosmer had 130 more PA than either of them. I have Ackley and Weeks both at 23 RAR with Hosmer at 21. Fielding and baserunning would seem to favor Weeks.
Greg Holland deserves a mention at least a mention as a reliever. Holland stranded 31 of 33 baserunners, the second-best performance of any AL reliever, and his peripherals were terrific as well. However, his 26 RAR is thanks in large part to the inherited runner performance, and thanks to Hosmer I wouldn’t be comfortable naming him the most valuable rookie on his own team. So I see it as:
1) SP Jeremy Hellickson, TB
2) SP Michael Pineda, SEA
3) SP Ivan Nova, NYA
4) 2B Jemile Weeks, OAK
5) 2B Dustin Ackley, SEA
You’ll note that I consider Mark Trumbo an afterthought. Yes, he hit 29 homers, but he also drew just 25 walks. His .290 OBA was second-lowest among AL first baseman with 300 PA, so despite the power, he ranks in the middle of the pack offensively at his position. He wouldn’t crack my top ten.
If Trumbo is the biggest source of divergence from my take on the award and the mainstream, his NL counterpart will certainly be Craig Kimbrel. Kimbrel was terrific by any measure, but in the end you have 77 innings pitched. I don’t believe in extreme leverage bonuses--or much of a leverage bonus at all. I’ll give him an arbitrary 25% boost to get to 25 RAR, but no more.
Among position players, the three standouts are Kimbrel’s teammate Freddie Freeman and Washington teammates Wilson Ramos and Danny Espinosa. I have them all essentially even in terms of RAR at 27. BP’s FRAA likes Espinosa’s fielding and baserunning, and that’s enough to put him in the lead. I suspect Freeman will get more support than Ramos, but the two aren’t that far apart as hitters, with Freeman creating 5.3 runs per game and Ramos 5.0. Freeman had nearly 200 more PA, but Ramos is a catcher. Freeman’s fielding reputation is good, but his FRAA was -5. It can go either way, but I prefer Ramos.
Josh Collmenter and Vance Worley were the top starters, with apologies to Cory Luebke, who I could certainly make a ballot case for, but will refrain lest I be accused of favoritism. Collmenter worked 23 more innings than Worley, which puts him 5 RAR ahead (36 to 31). Collmenter did have a BABIP of just .263 to Worley’s .293, but the dRA difference is not large enough (4.06 to 3.72) to convince me to put Worley ahead.
Depending on how you value Espinosa’s fielding, you certainly could conclude that he was more valuable than Collmenter--conservatively, I’ll stick with the later, and so my ballot is:
1) SP Josh Collmenter, ARI
2) 2B Danny Espinosa, WAS
3) SP Vance Worley, PHI
4) C Wilson Ramos, WAS
5) RP Craig Kimbrel, ATL
Monday, October 31, 2011
IBA Ballot: Rookie of the Year
Friday, October 28, 2011
Baseball
It has been a part of my life for almost as long as I can remember and it will remain so for as long as I live. For seven months of the year, it is as familiar a part of my life as brushing my teeth or eating dinner, and so it is easy to take for granted. But then one day I wake up and suddenly it is gone, and in the void there is malaise. When the weather is nice, it is played; when it is dark and cold, it moves towards the tropics and away from focus. While it can be used to tell seasons, it scoffs at time while it is played. The competitors dictate the endpoint through their play.
It is a team game, but in many ways it allows the individual to stand and be judged on his own merits. It is a game that, through its variants and offshoots, is quite playable by a large number of people. It is the great American pastime, but it is also the great Cuban passion, the great Dominican pastime, perhaps the most popular import Japan has ever known. We call it baseball, but it is equally beisbol, yakyu, honkbal, pelota.
It is a game simple enough that it can be described (and recorded, on nothing more complex than a piece of paper) discretely--by inning, by score, by out, by baserunner, by count--yet complex enough that there are hundreds and hundreds of people like me who are fascinated by it and spend much of our free time thinking about it, yet we still discover new things about it.
And if you are wired to view the world in a certain way, to try to find and verify patterns, to quantify when possible, and sometimes to find meaning and order through randomness and chance--then sabermetrics is a vessel for enjoying it, understanding it, and celebrating it. To know that what we have seen over the last month is not just unlikely--but rather to have a systematic way of thinking that allows us to estimate just how unlikely--does not detract from it.
Once in a while we are presented with just one more game--one game that is, without question, the end. It almost goes against the spirit of the game to be pettily constrained by a set limit of games that cannot be cheated, unlike the nine innings that often become ten, and sometimes become twelve, and on glorious occasions become twenty, and in theory can be infinite. The potential is often greater than the payoff--but either way, the journey was incredible.
Sunday, October 09, 2011
Brief Playoff Meanderings
* There have been eighteen postseasons in which the Division Series has been held (I’m counting the 1981 playoffs between the half-season winners as Division Series). 2011 set the new record for the most aggregate games played in the round, with nineteen. The maximum is twenty, and had the Rays managed to take an additional game from the Rangers it would have been reached. The previous high was eighteen, which occurred in 1981, 2001 and 2003.
The record for most total games played in the postseason (since 1995; in this case I’m excluding 1981 because the LCS was only a five-game series at that point) is 38 in 2003--two LDS went four and the World Series went six, but all other series went the distance. The ALCS and NLCS are both well-remembered (I can just say Grady Little or Aaron Boone and Steve Bartman and you’ll remember the circumstances).
No other postseason has come particularly close; the runner-up is 2001, which saw 35 total games played despite each LCS only lasting five games. The fewest games played in a post-season is 28 in 2007--every series was a sweep except for the two involving Cleveland, who beat New York in four in the ALDS then lost to Boston in seven in the ALCS. To put 2007 in perspective, every series from here on out in 2011 could be a sweep, and the total games played would be 31.
A natural follow-up question is “What is the expected number of postseason games?” If you assume that each game is a 50/50 proposition (equally matched teams, no home field advantage, no variation in team quality from day-to-day, etc.), then it’s very straightforward to estimate series length with the geometric distribution.
For a five-game series under those assumptions, there is a 25% chance for a sweep and a 37.5% chance for a four or five game series. For a seven-game series, there is a 12.5% chance for four games, 25% for five games, and 31.25% for six or seven games. Thus, the expected length of a five-game series is 4.125 games, the expected length of a seven-game series is 5.8125 games, and the expected number of games in the postseason is 33.9375. 1997, 2002 and 2004 all met expectations with 34 games.
However, if one compares the expected series lengths to the observed series length in the divisional era (1969 and foreward), he will find that five-game series do not conform to expectations:

Five-game series tend to be resolved in fewer games than one would expect assuming an equal probability of each outcome. The difference is statistically significant by reasonable standards. The average is just 3.86 games. Assuming that one of the teams has a .716 expected winning percentage comes close to minimizing the error assuming the geometric distribution framework:

I’m presenting this as a curiosity, and I’m certainly not suggesting that we should assume that the assumptions I described are useless when thinking about the Division Series. And on the other hand, seven-game series since 1969 conform almost as well as one could hope for:

There is a slight tendency for series to be resolved more quickly than one would expect, but it isn’t particularly significant, and the average of 5.75 is not far off the expected 5.81.
*What I’m going to say here is not in any way novel; many fans, both sabermetrically-inclined and not have expressed the same opinion over the years. But there were two instances that I considered so egregious in the Arizona/Milwaukee game give that I can’t help but comment on it here.
I have always thought that many managers are way too eager to make substitutions that sacrifice offense for baserunning or defense or the pitcher’s slot in the lineup, but I’m not sure I’ve ever seen a better display of it than in the aforementioned Game Five. In the eighth inning, Arizona trailed 2-1 with runners at first and third and two out. Chris Young drew a walk to load the bases and advance Miguel Montero from first to second, bringing Ryan Roberts up with the bases loaded.
At this point, Kirk Gibson decided to pinch-run for Montero, sending Collin Cowgill in. Montero occupied the #4 spot in the order, while Roberts was #7. Thus, it doesn’t take a rocket scientist to realize that, with an additional inning to go, there was a pretty good chance that Montero’s vacated spot would come up to bat again, and barring Arizona scoring at least two runs and holding Milwaukee in the bottom of the eighth, it would come with the Diamondbacks still needing a run (when I say needing a run, I mean it in the sense that Gibson apparently considered, since I would never say you don’t “need” more runs at any point in the game).
One would have to evaluate the marginal value of Cowgill’s baserunning very highly to see that as a winning move, especially considering that Montero would be off with contact given that their were two outs. Of course, as it played out, Roberts grounded into a fielder’s choice, and Montero’s spot did come up in the ninth, with the game now tied but runners at the corners and two outs. Henry Blanco hit into a fielder’s choice, and Arizona did not mount a threat in the tenth before allowing the game-winning run in the bottom of the frame.
The second move was not nearly as egregious, but it was still quite puzzling to me. With a 2-1 lead in the top of the ninth, Ron Roenicke summoned his closer, John Axford. The pitcher’s spot was due up fourth in the bottom of the ninth, so he double-switched Axford into Rickie Weeks’ #5 spot since he’d made the last out of the eighth.
Given that Roenicke wanted to make a double switch, Weeks was the only obvious candidate to be replaced--removing Braun or Fielder would be worse, especially since they were closer to coming to the plate, and Nyjer Morgan’s second spot was due up sixth in the bottom of the ninth. (One could make a case that Morgan would be the best candidate, but given that he got the walkoff hit in the tenth it wouldn’t be an argument that would fly over well with the “results not process” crowd).
What I find interesting about the double-switch for the home team taking the lead into the top of the ninth is that the only way the batting order matters at all is if Axford surrenders the lead. Thus, while you preserve Axford’s ability to pitch the tenth without sabotaging your offense in the ninth, you also know that if he does so, it will be only after he yielded a run in the ninth. You know that you will “need” runs if the #5 spot ever comes to the plate again.
Of course, this all worked out for Roenicke, since Axford pitched a 1-2-3 tenth, Morgan got the game-winning hit, and the #5 spot never batted again. And Roenicke does apparently like to bring Counsell in as a defensive replacement for Weeks, so if Weeks is going to come out of the game anyway, the double switch is the way to do it.
Sunday, October 02, 2011
End of Season Statistics 2011
The spreadsheets are published as Google Spreadsheets, which you can download in Excel format by changing the extension in the address from "=html" to "=xls". That way you can download them and manipulate things however you see fit.
The data comes from a number of different sources. Most of the basic data comes from Doug's Stats, which is a very handy site. KJOK's park database provided some of the data used in the park factors, but for recent seasons park data comes from anywhere that has it--Doug's Stats, or Baseball-Reference, or ESPN.com, or MLB.com. Data on pitcher's batted ball types allowed, doubles/triples allowed, and inherited/bequeathed runners comes from Baseball Prospectus.
The basic philosophy behind these stats is to use the simplest methods that have acceptable accuracy. Of course, "acceptable" is in the eye of the beholder, namely me. I use Pythagenpat not because other run/win converters, like a constant RPW or a fixed exponent are not accurate enough for this purpose, but because it's mine and it would be kind of odd if I didn't use it.
If I seem to be a stickler for purity in my critiques of others' methods, I'd contend it is usually in a theoretical sense, not an input sense. So when I exclude hit batters, I'm not saying that hit batters are worthless or that they *should* be ignored; it's just easier not to mess with them and not that much less accurate.
I also don't really have a problem with people using sub-standard methods (say, Basic RC) as long as they acknowledge that they are sub-standard. If someone pretends that Basic RC doesn't undervalue walks or cause problems when applied to extreme individuals, I'll call them on it; if they explain its shortcomings but use it regardless, I accept that. Take these last three paragraphs as my acknowledgment that some of the statistics displayed here have shortcomings as well.
The League spreadsheet is pretty straightforward--it includes league totals and averages for a number of categories, most or all of which are explained at appropriate junctures throughout this piece. The advent of interleague play has created two different sets of league totals--one for the offense of league teams and one for the defense of league teams. Before interleague play, these two were identical. I do not present both sets of totals (you can figure the defensive ones yourself from the team spreadsheet, if you desire), just those for the offenses. The exception is for the defense-specific statistics, like innings pitched and quality starts. The figures for those categories in the league report are for the defenses of the league's teams. However, I do include each league's breakdown of basic pitching stats between starters and relievers (denoted by "s" or "r" prefixes), and so summing those will yield the totals from the pitching side. The one abbreviation you might not recognize is "N"--this is the league average of runs/game for one team, and it will pop up again.
The Team spreadsheet focuses on overall team performance--wins, losses, runs scored, runs allowed. The columns included are: Park Factor (PF), Home Run Park Factor (PFhr), Winning Percentage (W%), Expected W% (EW%), Predicted W% (PW%), wins, losses, runs, runs allowed, Runs Created (RC), Runs Created Allowed (RCA), Home Winning Percentage (HW%), Road Winning Percentage (RW%) [exactly what they sound like--W% at home and on the road], Runs/Game (R/G), Runs Allowed/Game (RA/G), Runs Created/Game (RCG), Runs Created Allowed/Game (RCAG), and Runs Per Game (the average number of runs scored an allowed per game). Ideally, I would use outs as the denominator, but for teams, outs and games are so closely related that I don’t think it’s worth the extra effort.
The runs and Runs Created figures are unadjusted, but the per-game averages are park-adjusted, except for RPG which is also raw. Runs Created and Runs Created Allowed are both based on a simple Base Runs formula. The formula is:
A = H + W - HR - CS
B = (2TB - H - 4HR + .05W + 1.5SB)*.76
C = AB - H
D = HR
Naturally, A*B/(B + C) + D.
I have explained the methodology used to figure the PFs before, but the cliff’s notes version is that they are based on five years of data when applicable, include both runs scored and allowed, and they are regressed towards average (PF = 1), with the amount of regression varying based on the number of years of data used. There are factors for both runs and home runs. The initial PF (not shown) is:
iPF = (H*T/(R*(T - 1) + H) + 1)/2
where H = RPG in home games, R = RPG in road games, T = # teams in league (14 for AL and 16 for NL). Then the iPF is converted to the PF by taking x*iPF + (1-x), where x = .6 if one year of data is used, .7 for 2, .8 for 3, and .9 for 4+.
It is important to note, since there always seems to be confusion about this, that these park factors already incorporate the fact that the average player plays 50% on the road and 50% at home. That is what the adding one and dividing by 2 in the iPF is all about. So if I list Fenway Park with a 1.02 PF, that means that it actually increases RPG by 4%.
In the calculation of the PFs, I did not get picky and take out “home” games that were actually at neutral sites, like the Astros/Cubs series that was moved to Milwaukee in 2008.
There are also Team Offense and Defense spreadsheets. These include the following categories:
Team offense: Plate Appearances, Batting Average (BA), On Base Average (OBA), Slugging Average (SLG), Secondary Average (SEC), Walks per At Bat (WAB), Isolated Power (SLG - BA), R/G at home (hR/G), and R/G on the road (rR/G) BA, OBA, SLG, WAB, and ISO are park-adjusted by dividing by the square root of park factor (or the equivalent; WAB = (OBA - BA)/(1 - OBA) and ISO = SLG - BA).
Team defense: Innings Pitched, BA, OBA, SLG, Innings per Start (IP/S), Starter's eRA (seRA), Reliever's eRA (reRA), RA/G at home (hRA/G), RA/G on the road (rRA/G), Battery Mishap Rate (BMR), Modified Fielding Average (mFA), and Defensive Efficiency Record (DER). BA, OBA, and SLG are park-adjusted by dividing by the square root of PF; seRA and reRA are divided by PF.
The three fielding metrics I've included are limited it only to metrics that a) I can calculate myself and b) are based on the basic available data, not specialized PBP data. The three metrics are explained in this post, but here are quick descriptions of each:
1) BMR--wild pitches and passed balls per 100 baserunners = (WP + PB)/(H + W - HR)*100
2) mFA--fielding average removing strikeouts and assists = (PO - K)/(PO - K + E)
3) DER--the Bill James classic, using only the PA-based estimate of plays made. Based on a suggestion by Terpsfan101, I've tweaked the error coefficient. Plays Made = PA - K - H - W - HR - HB - .64E and DER = PM/(PM + H - HR + .64E)
Next are the individual player reports. I defined a starting pitcher as one with 15 or more starts. All other pitchers are eligible to be included as a reliever. If a pitcher has 40 appearances, then they are included. Additionally, if a pitcher has 50 innings and less than 50% of his appearances are starts, he is also included as a reliever (this allows some swingmen type pitchers who wouldn’t meet either the minimum start or appearance standards to get in).
For all of the player reports, ages are based on simply subtracting their year of birth from 2011. I realize that this is not compatible with how ages are usually listed and so “Age 27” doesn’t necessarily correspond to age 27 as I list it, but it makes everything a heckuva lot easier, and I am more interested in comparing the ages of the players to their contemporaries, for which case it makes very little difference. The "R" category records rookie status with a "R" for rookies and a blank for everyone else; I've trusted Baseball Prospectus on this. Also, all players are counted as being on the team with whom they played/pitched (IP or PA as appropriate) the most.
For relievers, the categories listed are: Games, Innings Pitched, Run Average (RA), Relief Run Average (RRA), Earned Run Average (ERA), Estimated Run Average (eRA), DIPS Run Average (dRA), Batted Ball Run Average (cRA), SIERA-style Run Average (sRA), Guess-Future (G-F), Inherited Runners per Game (IR/G), Batting Average on Balls in Play (%H), Runs Above Average (RAA), and Runs Above Replacement (RAR).
IR/G is per relief appearance (G - GS); it is an interesting thing to look at, I think, in lieu of actual leverage data. You can see which closers come in with runners on base, and which are used nearly exclusively to start innings. Of course, you can’t infer too much; there are bad relievers who come in with a lot of people on base, not because they are being used in high leverage situations, but because they are long men being used in low-leverage situations already out of hand.
For starting pitchers, the columns are: Wins, Losses, Innings Pitched, RA, RRA, ERA, eRA, dRA, cRA, sRA, G-F, %H, Pitches/Start (P/S), Quality Start Percentage (QS%), RAA, and RAR. RA and ERA you know--R*9/IP or ER*9/IP, park-adjusted by dividing by PF. The formulas for eRA, dRA, cRA, and sRA are in this article; I'm not going to copy them here, but all of them are based on the same Base Runs equation and they all estimate RA, not ERA:
* eRA is based on the actual results allowed by the pitcher (hits, doubles, home runs, walks, strikeouts, etc.). It is park-adjusted by dividing by PF.
* dRA is the classic DIPS-style RA, assuming that the pitcher allows a league average %H, and that his hits in play have a league-average S/D/T split. It is park-adjusted by dividing by PF.
* cRA is based on batted ball type (FB, GB, POP, LD) allowed, using the actual estimated linear weight value for each batted ball type. It is not park-adjusted.
* sRA is a SIERA-style RA, based on batted balls but broken down into just groundballs and non-groundballs. It is not park-adjusted either.
Both cRA and sRA are running a little high when compared to actual RA for 2010. Both measures are very sensitive and need to be recalibrated in order to overcome batted ball-type definition differences, frequencies of hit types on each kind of batted ball, and other factors, so keep in mind that they may not perfectly track RA without those adjustments (which I have not made in this case). I’ll let you make your own determination as to whether you find this data useful at all. Personally, I prefer to look at RRA, eRA, and dRA.
G-F is a junk stat, included here out of habit because I've been including it for years. It was intended to give a quick read of a pitcher's expected performance in the next season, based on eRA and strikeout rate. Although the numbers vaguely resemble RAs, it's actually unitless. As a rule of thumb, anything under four is pretty good for a starter. G-F = 4.46 + .095(eRA) - .113(K*9/IP). It is a junk stat. JUNK STAT JUNK STAT JUNK STAT. Got it?
%H is BABIP, more or less; I use an estimate of PA (IP*x + H + W, where x is the league average of (AB - H)/IP). %H = (H - HR)/(IP*x + H - HR - K). Pitches/Start includes all appearances, so I've counted relief appearances as one-half of a start (P/S = Pitches/(.5*G + .5*GS). QS% is just QS/(G - GS); I don't think it's particularly useful, but Doug's Stats include QS so I include it.
I've used a stat called Relief Run Average (RRA) in the past, based on Sky Andrecheck's article in the August 1999 By the Numbers; that one only used inherited runners, but I've revised it to include bequeathed runners as well, making it equally applicable to starters and relievers. I am using RRA as the building block for baselined value estimates for all pitchers this year. I explained RRA in this article, but the bottom line formulas are:
BRSV = BRS - BR*i*sqrt(PF)
IRSV = IR*i*sqrt(PF) - IRS
RRA = ((R - (BRSV + IRSV))*9/IP)/PF
The two baselined stats are Runs Above Average (RAA) and Runs Above Replacement (RAR). RAA uses the league average runs/game (N) for both starters and relievers, while RAR uses separate replacement levels for starters and relievers. Thus, RAA and RAR will be pretty close for relievers:
RAA = (N - RRA)*IP/9
RAR (relievers) = (1.11*N - RRA)*IP/9
RAR (starters) = (1.28*N - RRA)*IP/9
All players with 285 or more plate appearances are included in the Hitters spreadsheets. (I usually use 300 as a cutoff, but this year when I had the list sorted there were a number of players just below 300 that I was interested in, so I chose an arbitrarily lower threshold). Each is assigned one position, the one at which they appeared in the most games. The statistics presented are: Games played (G), Plate Appearances (PA), Outs (O), Batting Average (BA), On Base Average (OBA), Slugging Average (SLG), Secondary Average (SEC), Runs Created (RC), Runs Created per Game (RG), Speed Score (SS), Hitting Runs Above Average (HRAA), Runs Above Average (RAA), Hitting Runs Above Replacement (HRAR), and Runs Above Replacement (RAR).
I do not bother to include hit batters, so take note of that for players who do get plunked a lot. Therefore, PA are simply AB + W. Outs are AB - H + CS. BA and SLG you know, but remember that without HB and SF, OBA is just (H + W)/(AB + W). Secondary Average = (TB - H + W)/AB = SLG - BA + (OBA - BA)/(1 - OBA). I have not included net steals as many people (and Bill James himself) do--it is solely hitting events.
BA, OBA, and SLG are park-adjusted by dividing by the square root of PF. This is an approximation, of course, but I'm satisfied that it works well. The goal here is to adjust for the win value of offensive events, not to quantify the exact park effect on the given rate. I use the BA/OBA/SLG-based formula to figure SEC, so it is park-adjusted as well.
Runs Created is actually Paul Johnson's ERP, more or less. Ideally, I would use a custom linear weights formula for the given league, but ERP is just so darn simple and close to the mark that it’s hard to pass up. I still use the term “RC” partially as a homage to Bill James (seriously, I really like and respect him even if I’ve said negative things about RC and Win Shares), and also because it is just a good term. I like the thought put in your head when you hear “creating” a run better than “producing”, “manufacturing”, “generating”, etc. to say nothing of names like “equivalent” or “extrapolated” runs. None of that is said to put down the creators of those methods--there just aren’t a lot of good, unique names available. Anyway, RC = (TB + .8H + W + .7SB - CS - .3AB)*.322.
RC is park adjusted by dividing by PF, making all of the value stats that follow park adjusted as well. RG, the rate, is RC/O*25.5. I do not believe that outs are the proper denominator for an individual rate stat, but I also do not believe that the distortions caused are that bad. (I still intend to finish my rate stat series and discuss all of the options in excruciating detail, but alas you’ll have to take my word for it now).
I have decided to switch to a watered-down version of Bill James' Speed Score this year; I only use four of his categories. Previously I used my own knockoff version called Speed Unit, but trying to keep it from breaking down every few years was a wasted effort.
Speed Score is the average of four components, which I'll call a, b, c, and d:
a = ((SB + 3)/(SB + CS + 7) - .4)*20
b = sqrt((SB + CS)/(S + W))*14.3
c = ((R - HR)/(H + W - HR) - .1)*25
d = T/(AB - HR - K)*450
James actually uses a sliding scale for the triples component, but it strikes me as needlessly complex and so I've streamlined it. I also changed some of his division to mathematically equivalent multiplications.
There are a whopping four categories that compare to a baseline; two for average, two for replacement. Hitting RAA compares to a league average hitter; it is in the vein of Pete Palmer’s Batting Runs. RAA compares to an average hitter at the player’s primary position. Hitting RAR compares to a “replacement level” hitter; RAR compares to a replacement level hitter at the player’s primary position. The formulas are:
HRAA = (RG - N)*O/25.5
RAA = (RG - N*PADJ)*O/25.5
HRAR = (RG - .73*N)*O/25.5
RAR = (RG - .73*N*PADJ)*O/25.5
PADJ is the position adjustment, and it is based on 1992-2001 offensive data. For catchers it is .89; for 1B/DH, 1.19; for 2B, .93; for 3B, 1.01; for SS, .86; for LF/RF, 1.12; and for CF, 1.02. It dawned on me when re-reading this before posting that the timeframe means that I’ve been using the same PADJ for ten years--which means two things:
1) I’m getting old
2) It’s probably time for an update. I’ll look at 2002-2011 in my forthcoming annual “Offense by Postion” post
That was the mechanics of the calculations; now I'll twist myself into knots trying to justify them. If you only care about the how and not the why, stop reading now.
The first thing that should be covered is the philosophical position behind the statistics posted here. They fall on the continuum of ability and value in what I have called "performance". Performance is a technical-sounding way of saying "Whatever arbitrary combination of ability and value I prefer".
With respect to park adjustments, I am not interested in how any particular player is affected, so there is no separate adjustment for lefties and righties for instance. The park factor is an attempt to determine how the park affects run scoring rates, and thus the win value of runs.
I apply the park factor directly to the player's statistics, but it could also be applied to the league context. The advantage to doing it my way is that it allows you to compare the component statistics (like Runs Created or OBA) on a park-adjusted basis. The drawback is that it creates a new theoretical universe, one in which all parks are equal, rather than leaving the player grounded in the actual context in which he played and evaluating how that context (and not the player's statistics) was altered by the park.
The good news is that the two approaches are essentially equivalent; in fact, they are equivalent if you assume that the Runs Per Win factor is equal to the RPG. Suppose that we have a player in an extreme park (PF = 1.15, approximately like Coors Field pre-humidor) who has an 8 RG before adjusting for park, while making 350 outs in a 4.5 N league. The first method of park adjustment, the one I use, converts his value into a neutral park, so his RG is now 8/1.15 = 6.957. We can now compare him directly to the league average:
RAA = (6.957 - 4.5)*350/25.5 = +33.72
The second method would be to adjust the league context. If N = 4.5, then the average player in this park will create 4.5*1.15 = 5.175 runs. Now, to figure RAA, we can use the unadjusted RG of 8:
RAA = (8 - 5.175)*350/25.5 = +38.77
These are not the same, as you can obviously see. The reason for this is that they take place in two different contexts. The first figure is in a 9 RPG (2*4.5) context; the second figure is in a 10.35 RPG (2*4.5*1.15) context. Runs have different values in different contexts; that is why we have RPW converters in the first place. If we convert to WAA (using RPW = RPG), then we have:
WAA = 33.72/9 = +3.75
WAA = 38.77/10.35 = +3.75
Once you convert to wins, the two approaches are equivalent. The other nice thing about the first approach is that once you park-adjust, everyone in the league is in the same context, and you can dispense with the need for converting to wins at all. You still might want to convert to wins, and you'll need to do so if you are comparing the 2010 players to players from other league-seasons (including between the AL and NL in the same year), but if you are only looking to compare Jose Bautista to Miguel Cabrera, it's not necessary. WAR is somewhat ubiquitous now, but personally I prefer runs when possible--why mess with decimal points if you don't have to?
The park factors used to adjust player stats here are run-based. Thus, they make no effort to project what a player "would have done" in a neutral park, or account for the difference effects parks have on specific events (walks, home runs, BA) or types of players. They simply account for the difference in run environment that is caused by the park (as best I can measure it). As such, they don't evaluate a player within the actual run context of his team's games; they attempt to restate the player's performance as an equivalent performance in a neutral park.
I suppose I should also justify the use of sqrt(PF) for adjusting component statistics. The classic defense given for this approach relies on basic Runs Created--runs are proportional to OBA*SLG, and OBA*SLG/PF = OBA/sqrt(PF)*SLG/sqrt(PF). While RC may be an antiquated tool, you will find that the square root adjustment is fairly compatible with linear weights or Base Runs as well. I am not going to take the space to demonstrate this claim here, but I will some time in the future.
Many value figures published around the sabersphere adjust for the difference in quality level between the AL and NL. I don't, but this is a thorny area where there is no right or wrong answer as far as I'm concerned. I also do not make an adjustment in the league averages for the fact that the overall NL averages include pitcher batting and the AL does not (not quite true in the era of interleague play, but you get my drift).
The difference between the leagues may not be precisely calculable, and it certainly is not constant, but it is real. If the average player in the AL is better than the average player in the NL, it is perfectly reasonable to expect the average AL player to have more RAR than the average NL player, and that will not happen without some type of adjustment. On the other hand, if you are only interested in evaluating a player relative to his own league, such an adjustment is not necessarily welcome.
The league argument only applies cleanly to metrics baselined to average. Since replacement level compares the given player to a theoretical player that can be acquired on the cheap, the same pool of potential replacement players should by definition be available to the teams of each league. One could argue that if the two leagues don't have equal talent at the major league level, they might not have equal access to replacement level talent--except such an argument is at odds with the notion that replacement level represents talent that is truly "freely available".
So it's hard to justify the approach I take, which is to set replacement level relative to the average runs scored in each league, with no adjustment for the difference in the leagues. The best justification is that it's simple and it treats each league as its own universe, even if in reality they are connected.
The replacement levels I have used here are very much in line with the values used by other sabermetricians. This is based both on my own "research", my interpretation of other's people research, and a desire to not stray from consensus and make the values unhelpful to the majority of people who may encounter them.
Replacement level is certainly not settled science. There is always going to be room to disagree on what the baseline should be. Even if you agree it should be "replacement level", any estimate of where it should be set is just that--an estimate. Average is clean and fairly straightforward, even if its utility is questionable; replacement level is inherently messy. So I offer the average baseline as well.
For position players, replacement level is set at 73% of the positional average RG (since there's a history of discussing replacement level in terms of winning percentages, this is roughly equivalent to .350). For starting pitchers, it is set at 128% of the league average RA (.380), and for relievers it is set at 111% (.450).
I am still using an analytical structure that makes the comparison to replacement level for a position player by applying it to his hitting statistics. This is the approach taken by Keith Woolner in VORP (and some other earlier replacement level implementations), but the newer metrics (among them Rally and Fangraphs' WAR) handle replacement level by subtracting a set number of runs from the player's total runs above average in a number of different areas (batting, fielding, baserunning, positional value, etc.), which for lack of a better term I will call the subtraction approach.
The offensive positional adjustment makes the inherent assumption that the average player at each position is equally valuable. I think that this is close to being true, but it is not quite true. The ideal approach would be to use a defensive positional adjustment, since the real difference between a first baseman and a shortstop is their defensive value. When you bat, all runs count the same, whether you create them as a first baseman or as a shortstop.
That being said, using “replacement hitter at position” does not cause too many distortions. It is not theoretically correct, but it is practically powerful. For one thing, most players, even those at key defensive positions, are chosen first and foremost for their offense. Empirical work by Keith Woolner has shown that the replacement level hitting performance is about the same for every position, relative to the positional average.
Figuring what the defensive positional adjustment should be, though, is easier said than done. Therefore, I use the offensive positional adjustment. So if you want to criticize that choice, or criticize the numbers that result, be my guest. But do not claim that I am holding this up as the correct analytical structure. I am holding it up as the most simple and straightforward structure that conforms to reality reasonably well, and because while the numbers may be flawed, they are at least based on an objective formula that I can figure myself. If you feel comfortable with some other assumptions, please feel free to ignore mine.
That still does not justify the use of HRAR--hitting runs above replacement--which compares each hitter, regardless of position, to 73% of the league average. Basically, this is just a way to give an overall measure of offensive production without regard for position with a low baseline. It doesn't have any real baseball meaning.
A player who creates runs at 90% of the league average could be above-average (if he's a shortstop or catcher, or a great fielder at a less important fielding position), or sub-replacement level (DHs that create 4 runs a game are not valuable properties). Every player is chosen because his total value, both hitting and fielding, is sufficient to justify his inclusion on the team. HRAR fails even if you try to justify it with a thought experiment about a world in which defense doesn't matter, because in that case the absolute replacement level (in terms of RG, without accounting for the league average) would be much higher than it is currently.
The specific positional adjustments I use are based on 1992-2001 data. There's no particular reason for not updating them; at the time I started using them, they represented the ten most recent years. I have stuck with them because I have not seen compelling evidence of a change in the degree of difficulty or scarcity between the positions between now and then, and because I think they are fairly reasonable. The positions for which they diverge the most from the defensive position adjustments in common use are 2B, 3B, and CF. Second base is considered a premium position by the offensive PADJ (.94), while third base and center field are both neutral (1.01 and 1.02).
Another flaw is that the PADJ is applied to the overall league average RG, which is artificially low for the NL because of pitcher's batting. When using the actual league average runs/game, it's tough to just remove pitchers--any adjustment would be an estimate. If you use the league total of runs created instead, it is a much easier fix.
One other note on this topic is that since the offensive PADJ is a proxy for average defensive value by position, ideally it would be applied by tying it to defensive playing time. I have done it by outs, though.
The reason I have taken this flawed path is because 1) it ties the position adjustment directly into the RAR formula rather then leaving it as something to subtract on the outside and more importantly 2) there’s no straightforward way to do it. The best would be to use defensive innings--set the full-time player to X defensive innings, figure how Derek Jeter’s innings compared to X, and adjust his PADJ accordingly. Games in the field or games played are dicey because they can cause distortion for defensive replacements. Plate Appearances avoid the problem that outs have of being highly related to player quality, but they still carry the illogic of basing it on offensive playing time. And of course the differences here are going to be fairly small (a few runs). That is not to say that this way is preferable, but it’s not horrible either, at least as far as I can tell.
To compare this approach to the subtraction approach, start by assuming that a replacement level shortstop would create .86*.73*4.5 = 2.825 RG (or would perform at an overall level of equivalent value to being an average fielder at shortstop while creating 2.825 runs per game). Suppose that we are comparing two shortstops, each of whom compiled 600 PA and played an equal number of defensive games and innings (and thus would have the same positional adjustment using the subtraction approach). Alpha made 380 outs and Bravo made 410 outs, and each ranked as dead-on average in the field.
The difference in overall RAR between the two using the subtraction approach would be equal to the difference between their offensive RAA compared to the league average. Assuming the league average is 4.5 runs, and that both Alpha and Bravo created 75 runs, their offensive RAAs are:
Alpha = (75*25.5/380 - 4.5)*380/25.5 = +7.94
Similarly, Bravo is at +2.65, and so the difference between them will be 5.29 RAR.
Using the flawed approach, Alpha's RAR will be:
(75*25.5/380 - 4.5*.73*.86)*380/25.5 = +32.90
Bravo's RAR will be +29.58, a difference of 3.32 RAR, which is two runs off of the difference using the subtraction approach.
The downside to using PA is that you really need to consider park effects if you, whereas outs allow you to sidestep park effects. Outs are constant; plate appearances are linked to OBA. Thus, they not only depend on the offensive context (including park factor), but also on the quality of one's team. Of course, attempting to adjust for team PA differences opens a huge can of worms which is not really relevant; for now, the point is that using outs for individual players causes distortions, sometimes trivial and sometimes bothersome, but almost always makes one's life easier.
I do not include fielding (or baserunning outside of steals, although that is a trivial consideration in comparison) in the RAR figures--they cover offense and positional value only). This in no way means that I do not believe that fielding is an important consideration in player valuation. However, two of the key principles of these stat reports are 1) not incorporating any data that is not readily available and 2) not simply including other people's results (of course I borrow heavily from other people's methods, but only adapting methodology that I can apply myself).
Any fielding metric worth its salt will fail to meet either criterion--they use zone data or play-by-play data which I do not have easy access to. I do not have a fielding metric that I have stapled together myself, and so I would have to simply lift other analysts' figures.
Setting the practical reason for not including fielding aside, I do have some reservations about lumping fielding and hitting value together in one number because of the obvious differences in reliability between offensive and fielding metrics. In theory, they absolutely should be put together. But in practice, I believe it would be better to regress the fielding metric to a point at which it would be roughly equivalent in reliability to the offensive metric.
Offensive metrics have error bars associated with them, too, of course, and in evaluating a single season's value, I don't care about the vagaries that we often lump together as "luck". Still, there are errors in our assessment of linear weight values and players that collect an unusual proportion of infield hits or hits to the left side, errors in estimation of park factor, and any number of other factors that make their events more or less valuable than an average event of that type.
Fielding metrics offer up all of that and more, as we cannot be nearly as certain of true successes and failures as we are when analyzing offense. Recent investigations, particularly by Colin Wyers, have raised even more questions about the level of uncertainty. So, even if I was including a fielding value, my approach would be to assume that the offensive value was 100% reliable (which it isn't), and regress the fielding metric relative to that (so if the offensive metric was actually 70% reliable, and the fielding metric 40% reliable, I'd treat the fielding metric as .4/.7 = 57% reliable when tacking it on, to illustrate with a simplified and completely made up example presuming that one could have a precise estimate of nebulous "reliability").
Given the inherent assumption of the offensive PADJ that all positions are equally valuable, once RAR has been figured for a player, fielding value can be accounted for by adding on his runs above average relative to a player at his own position. If there is a shortstop that is -2 runs defensively versus an average shortstop, he is without a doubt a plus defensive player, and a more valuable defensive player than a first baseman who was +1 run better than an average first baseman. Regardless, since it was implicitly assumed that they are both average defensively for their position when RAR was calculated, the shortstop will see his value docked two runs. This DOES NOT MEAN that the shortstop has been penalized for his defense. The whole process of accounting for positional differences, going from hitting RAR to positional RAR, has benefited him.
I've found that there is often confusion about the treatment of first baseman and designated hitters in my PADJ methodology, since I consider DHs as in the same pool as first baseman. The fact of the matter is that first baseman outhit DH. There is any number of potential explanations for this; DHs are often old or injured, players hit worse when DHing than they do when playing the field, etc. This actually helps first baseman, since the DHs drag the average production of the pool down, thus resulting in a lower replacement level than I would get if I considered first baseman alone.
However, this method does assume that a 1B and a DH have equal defensive value. Obviously, a DH has no defensive value. What I advocate to correct this is to treat a DH as a bad defensive first baseman, and thus knock another five or ten runs off of his RAR for a full-time player. I do not incorporate this into the published numbers, but you should keep it in mind. However, there is no need to adjust the figures for first baseman upwards --the only necessary adjustment is to take the DHs down a notch.
Finally, I consider each player at his primary defensive position (defined as where he appears in the most games), and do not weight the PADJ by playing time. This does shortchange a player like Ben Zobrist (who saw significant time at a tougher position than his primary position), and unduly boost a player like Buster Posey (who logged a lot of games at a much easier position than his primary position). For most players, though, it doesn't matter much. I find it preferable to make manual adjustments for the unusual cases rather than add another layer of complexity to the whole endeavor.
Player spreadsheets should be coming by the middle of the week.
2011 Park Factors
2011 Leagues
2011 Teams
2011 Team Offense
2011 Team Defense
2011 AL Relievers
2011 NL Relievers
2011 AL Starters
2011 NL Starters
2011 AL Hitters
2011 NL Hitters
Thursday, September 29, 2011
Playoff Meanderings
I always like to put down some of my thoughts about the playoffs each year, but it’s a challenge to say anything even remotely close to being meaningful. Predicting the outcome of short series is folly (although I’ll engage in a little of this folly later), and you can read that anywhere. So I always try to come up with a different angle to illustrate why the playoffs are subject to such uncertainty.
I’ve certainly had some more interesting illustrations in the past; this one is pretty lame, but for some reason when it crossed my mind in August I thought it was a lot more interesting than I do now. What is the value of each playoff game or series in terms of a regular season game? In asking this, I’m not talking about the weight that should be applied to playoff performance for evaluating individual value, or any such thing...I’m just asking what the implied value is, given the assumption that the regular season standings carry over to the playoffs.
Of course, that’s not how it works--it's a cliché, but every team starts every series out at 0-0. However, I’ll assume that regular season standings carry over (resetting with the start of each additional round to keep things manageable) and the playoff games are weighted in a manner such that at the end of the series, the team that wins the series has a better overall record than its opponent.
There are at least two different ways to approach this increasingly silly scenario, which will be best illustrated by example--treating the series outcome as a binary, or considering the games individually. Suppose that the Alphas enter a five-game division series with a record of 92-70 while their opponents the Betas are 90-72.
First, from the series outcome perspective, if the Alphas win, the series was unnecessary since the Alphas already led in the standings. If the Betas win, however, the series must be given a weight of a number of games such that adding that many wins to the Betas and losses to the Alphas give the Betas a better record. Leaving things in terms of whole games, the answer in this case is three. Giving the Betas three additional wins leaves them at 93-72; three additional losses for the Alphas would make them 92-73. The series could have gone three, four, or five games, making the effective value of those games equal to either 1, .75, or .6 regular season games.
You can also consider this from the game perspective, that is actually looking at the outcome of each game in the series rather than treating the series as a binary win or loss. If the Betas win the above series 3-0, this is pretty straightforward given the two game margin--treating playoff games as equivalent to regular season games leaves the Alphas 92-73 and the Betas 93-72. Suppose the Betas had been 88-74 instead of 90-72, though. In order to bring the Betas ahead of the Alphas (on a whole wins basis), they need five, so each win (and thus each game has to be worth) 5/3 = 1.67 times a regular season game. Now the Alphas have 3*1.67 + 70 = 75 losses and the Betas have 3*1.67 + 88 = 93 wins, so that the Alphas record is 92-75 and the Betas 93-74.
You can see that if the final margin of the series is 3-2 in favor of the Betas, the weight on each playoff game would have to be roughly four times that of a regular season game since the Betas only pick up one win when the playoff series is considered. A 4x weight brings the Alphas and Betas together at 100-82.
This is all just a silly digression, but given the assumptions it is a simple way to think about how the implied value of a playoff game compares to that of a regular season game.
Getting to the 2011 playoffs, let me offer some quick thoughts. I’ll leave the detailed handicapping to those who are better suited for it and also like quixotic quests. The marginal value of more in-depth analysis is limited, but if that’s what you seek, you won’t find it here.
The probabilities that follow assume nothing about home field advantage or pitching matchups, or even true talent for that matter. They are simply based on my crude team rankings, fueled by 25% actual W%, 25% expected W% (from R/RA), 25% predicted W% (from RC/RCA), and 25% from .500.
That formula is also arbitrary. The results should be fairly reasonable, but I’m also eager to disown at the same time, as something of a commentary on the futileness of the exercise...and most especially the bloviating that is done without any logic at all. I’m sure that there are many scribes across the country furiously writing about how certain teams have no chance, never learning the lesson that the differences between major league teams simply aren’t that great, especially after eight of the best have been selected from a 162 game sample.

This method considers all of the playoff teams to be in the top ten in MLB; only Boston (#3) and the Angeles (#8) are on the outside looking in. The Yankees, Phillies, and Rangers are near co-favorites to win it all; NYA and TEX are ranked about evenly, while PHI benefits from the weaker NL field and has the highest odds of winning a first round series and the pennant. Overall, the AL has an estimated 57% chance of winning the World Series. The most likely matchup in the Series is NYA/PHI (11%); the least likely is DET/STL (4%). The rankings imply that the worst playoff team (ARI) would beat the best playoff team (NYA) 43% of the time, which over 162 games is seventy wins. Strictly equating true probability to actual 2011 record, consider the odds that the Padres could win a seven game series against the Indians, and there is roughly the same likelihood of the Diamondbacks winning a seven game series against the Yankees.
As far as my personal rooting interests go, New York and Tampa Bay are my top two choices, followed by Milwaukee and St. Louis. I would be happy to see any of those teams win, have no particularly strong feelings about Arizona or Detroit, and be mildly disappointed if it’s Philadelphia or Texas. But there are no White Sox in this group.
Tuesday, September 20, 2011
A Quick Look at Negro League W-L Records
I wrote this about a year ago and wasn’t sure if I’d ever post it. With the recent publication of some Negro League data at Seamheads, I figured I’d better post it now before it became completely dated. The data I used was compiled by Chris Cobb and posted on the Hall of Merit site, with John Holway's research as his source data.
I need to admit upfront that I know very little about the Negro Leagues. My knowledge level of the Negro Leagues peaked at about age eleven when I read Only the Ball Was White, and has only gone downhill since then. That is one of the reasons for this post--as a (very limited) education for me on the great pitchers of the Negro Leagues.
I am going to be applying the Netural Win-Loss record approach introduced by Rob Wood, which I have written about several times. It is a way to contextualize a pitcher's W-L record using only the win-loss record of the pitcher's team. This post applies it to several Negro League pitchers.
The basic idea behind Wood's approach is that an average team's deviation from .500 is due in equal parts to their offense and defense. The portion of a team's deviation from .500 that arises from the defense (with the exception of relievers in the pitcher's game and fielders) doesn't do anything to increase a pitcher's expected W% in reality, but if you compare his W% directly to that of his teammates', he will suffer for it.
The formula is simple and linear; instead of comparing a pitcher to his team's W% when he does not get a decision (Mate), the comparison is to the average of Mate and .500. The neutral W% is easy to figure:
NW% = W% - Mate/2 + .25
From NW%, one can figure Neutral Wins and Losses:
NW = NW%*(W + L), NL = W + L - NW
It is also very easy to combine NW% and the number of decisions into wins above some baseline. Wins Above Team is traditionally defined as wins above .500:
WAT = (NW% - .5)*(W + L)
I also use Wins Compared to Replacement, with the assumption that a replacement level starter will have a .380 W%:
WCR = (NW% - .38)*(W + L)
There are a number of weaknesses to the Neutral W-L approach, and there are a number of additional complications that arise when applying it to the Negro League data. This is an incomplete list of the methodological issues that are present even when looking at major league data:
* It does not isolate performance when the pitcher actually pitches; some will receive lousy run support despite pitching for good offensive teams.
* While the approach assumes that the team is balanced between offense and defense, this is not always the case. It is a decent assumption for a pitcher's entire career, but there are still going to be cases in which a pitcher is predominantly on teams skewed one way or the other. Those on offensive teams will benefit unfairly in the metric, while those who are on teams with otherwise strong starting pitching staffs will be hurt.
* All of the problems with the definition and concept behind pitcher wins and losses themselves are still present
With respect to the Negro League results included in this post, the data I have used was compiled by Chris Cobb and posted on the Hall of Merit site, with John Holway's research as his source data. Among the problems that arise from the data:
* The records themselves are incomplete (missing seasons, team records only published for half seasons, etc.) and sometimes contradictory (individual totals that don't add up to the team total, etc.) These kind of errors exist even in major league data from the period, so it's no surprise that they are present in the more chaotic, less-organized Negro League data.
To deal with the gaps in the specific data I used, if I couldn't find the team's record, I assumed that they were .500 when the pitcher in question's decisions were removed. If a pitcher split time between teams and there was no breakdown of his W-L record with the two teams provided, I used the average of the two team's record. For seasons in which Cobb did not include the team's record and I had to look it up from another source, I used the ESPN Baseball Encyclopedia. In that case, if the team's record was only available for a half-season, I assumed that the full season record was double the half-season record.
* I only used the results from domestic Negro League games. The world of the Negro Leagues encompassed a lot more than that; players went to the Caribbean to play, teams barnstormed extensively, played games against major league opponents, etc. Limiting the analysis to league games makes it workable, but it does omit a lot of relevant performances.
In this regard the Negro Leagues were similar to the early NA/NL days, in which the league schedule constituted only a small fraction of total games played, and independent teams often compared favorably to league opponents.
* I am way out of my area of knowledge, but even I feel comfortable asserting that the NeL pitching rotations looked more like the early majors then the contemporary majors. Pitchers got a higher percentage of their team's decisions, reducing the sample size from which Mate is drawn and weakening the assumption that the other pitchers are average. I have also read that teams would purposefully match their aces against one another to create gate attractions, whereas our normal assumption is that teams will try to match their pitchers up in whatever manner creates the highest number of expected wins.
* The league structure was less stable from year-to-year, which makes it harder to compare NeL pitchers from one time period to the other. For twentieth-century major league pitchers, we can be confident that, regardless of when they pitched, that they were facing the highest level of competition available (with the obvious exception of the players locked out of the majors due to their skin color). We also know that they pitched in seasons of roughly equal length, and so their career records represent a fair sample of their performance at different ages.
We don't have that confidence when dealing with the NeL data. For example, Satchel Paige gets no credit for 1935 here, but the adjacent seasons of 1934 and 1936 appear to be among his best. Then he gets no credit for 1937-39, as he was not pitching in official league games. You will see that Paige doesn't come out as impressively as might be expected in the career totals, but the gaps in league play might well be the major cause.
* I have listed WCR figures using a .380 replacement level, but in actuality I have no idea where the NeL replacement level should be set.
From all of the caveats, it may seem as if I am declaring the NW-L statistics to be useless. That is not my intention; I simply don't want to oversell them or fail to acknowledge their biases. Many of the issues with the NW-L records are issues that would arise with any statistical analysis of Negro League pitchers. Consider what a logistical nightmare it would be to try to look at runs allowed, needing innings, and league averages, and park factors.
As sabermetricians we all know the flaws of pitcher W-L records, but there are a few benefits. Among them is the ease in determining them, at least if complete games are the norm. All you need to know is who the starting pitcher was and which team won the game, and you've got it. No need for box scores or play-by-play. No need for park factors or league averages--the average in every league and every park for all of time is .500.
These useful properties are most useful when dealing with incomplete data, and we can refine them further by incorporating team record and producing NW-L. Are the results perfect? Absolutely not. Are they likely to give us a better indication of the quality of these pitchers than raw W-L record or uncontextualized ERAs? I say yes.
The pitchers for whom data was available were: Chet Brewer, Dave Brown, Ray Brown, Bill Byrd, Andy Cooper, Leon Day, Willie Foster, Leroy Matlock, Satchel Paige, Dick Redding, Bullet Joe Rogan, Hilton Smith, Smokey Joe Williams, and Nip Winters.
Since I am out of my area of knowledge when discussing the Negro League stars, I'm not going to make a lot of comments--I'll leave interpretation up to the reader. Here are the actual career W-L records for the pitchers, along with Mate. The list is sorted by career wins above .380:

Only one of the pitchers had a worse record than that of his teams (Chet Brewer). If one figures Wins Above Team by the traditional method, Brewer would rate as a below-average pitcher. It's far more likely, though, that a pitcher with a .591 W% regarded as an excellent pitcher was in fact an excellent pitcher. The fact that his teams played .624 baseball without him indicates that they probably had above-average pitching, which while good for the team did absolutely nothing to increase Brewer's expected W%. Brewer still takes a hit, of course, when neutralizing his record by the Wood approach, but is assumed to be an above-average performer.
Here are the career neutral W-L records for the pitchers, sorted by WCR:

Here is a link to the spreadsheet containing the complete yearly breakdowns for each pitcher. You can see exactly what I inputted and which seasons I didn't have team records for (you'll see blanks in the TW and TL columns):
https://docs.google.com/spreadsheet/pub?hl=en_US&hl=en_US&key=0AnPJbQnlHhRHdEl5TjRzUEVscjNQVy1naDY0ODVtZlE&output=html
Again, this is obviously a very incomplete examination of the careers of a limited number of Negro League stars, and I certainly would not advocate placing too much stock in the results.
Monday, September 12, 2011
Scoring Self-Indulgence, pt. 5: Baserunner Advances
Last time I covered my scoring codes to recognize a batter reaching base; this time I’ll discuss what I record once he gets there. I’ll start with advances made by the runner independent of the actions of a subsequent batter on his team--things like stolen bases and advancing on wild pitches. Most of the codes that follow are pretty straightforward. In each case, I’ll show the advance as a runner going from first to second, but the same concepts apply to advancing to third and scoring. In each case, I’ll assume that the batter reached first by being hit by a pitch.
For every advancement that occurs during the course of a plate appearance, I record both the lineup slot of the batter at the plate and the pitch on which the event occurs (or which pitches it is between if applicable). The pitch is indicated by the same letter used in the batter’s box, except in lower case--the first pitch of a PA is “a”, the second pitch is “b”, etc.
There are several exceptions. If the last pitch (which is never given a letter in the batter’s scorebox) is labeled “lp”. If an event happens before the first pitch of a plate appearance, I use “bfp”. Finally, if an event occurs between pitches, it is labeled “a!”, where ! is replaced by the pitch letter for the last pitch before the event. Suppose the event occurs between pitches two and three of the plate appearance; in this case, the pitch code for the event is “ab”, because the second pitch (b) was the last one thrown. “ab” can be read as “after b”.
The code for a stolen base is the obvious “SB”. If it occurred on the third pitch of a plate appearance taken by the #6 batter in the lineup, the scoring would be:

As you can see from the example, the pitch information is written above the advancement symbol, in smaller type.
Wild pitches and passed balls are separated by a distinction I’d wipe out of the rule book if given the chance, but I do record them differently: “WP” and “PB” are the obvious codes. In the examples, the wild pitch occurs on the last pitch to the #6 batter, and the passed ball occurs on the first pitch to the #2 hitter:


The code for a balk is “BK”; this one comes between the third and fourth pitches to the cleanup hitter:

I don’t like the scoring distinction between a stolen base and defensive indifference, but I do make note of it on my scoresheets because SB is such a common category, and it’s easier to keep track of the ones that are really scored as steals and add in fielder’s indifference if one chooses than it is to try to divine after the fact what the scoring was. I refer to it as Fielder’s Indifference (FI), because it is a subset of fielder’s choice by definition, and the symbol seems more consistent. This one occurs on the sixth pitch to the #5 hitter:

A runner could advance on an error between pitches, which almost always would be a throwing error. If the pitcher throws the ball away on a pickoff attempt before the first pitch to the #2 hitter, the scoring looks like this:

Sometimes, the extra bases are gained before the batter-runner becomes a runner; that is, on the same play on which he reaches base. Suppose a batter dribbles a hit to the pitcher, but in his haste to make the play, the pitcher hurls the ball down the right field line, allowing the batter to move up to second. The scoring looks like this:

If there is no additional information included with the notation, it is assumed that the advance occurred on the same play as the on-base event. The other common way a batter-runner moves up is when he is able to advance on a throw to another base made in an attempt to retire another runner. In this example, the batter singles to right, then advances to second on a throw home. The code “ATx” means advanced on throw, with x standing in for the base to which the throw is made (2 for second, 3 for third, and H for home):

I have yet to touch on the means by which most bases are gained: advances on plays initiated by subsequent batters. I mark these by writing and circling the batting order position of the batter responsible in the quadrant of the runner’s scorebox corresponding to the base he wound up at. Suppose the runner from first advances to third on a play initiated by the #7 hitter. I would score it:

If the runner scores, then I use a box instead of a circle, so that it’s easy to distinguish how many runs a team has scored. In this case, the runner who moved to third a player initiated by the #7 hitter ends up scoring on a play initiated by the #9 hitter:

A runner can also score due to an event not initiated by another batter. The most common is scoring on a wild pitch. In this example, the runner from third scores on a wild pitch, with the wild pitch coming on the second pitch to the #9 hitter:

As you can see, I allow the box that indicates a run scored to vary in shape and size as appropriate to allow the necessary space for recording the event.
Sometimes, the event that advances the baserunner occurs while the ball is in play, but referring to the relevant batter’s scorebox will not note how the advancement occurred. Suppose that there is a runner on first, and the batter singles to right, advancing the runner to second. Then the right fielder boots the ball, allowing both to move up one base (the batter-runner to second and the runner to third). In this case, I would simply record the appropriate batter number circled in the runner’s third base quadrant. The batter’s scorebox will include the error by the right fielder, imply that it occurred during his plate appearances, and thus imply that the single plus the error enabled the runner to advance from first to third.
However, there are also cases in which the runner advances but the batter stays put. Suppose that the same play occurs as described above, except the batter-runner (who happens to be the #4 hitter) stops at first base. Now I would score the runner’s advancement as:

In this case, the use of a small circled 4 above the error indicates that the error occurred during the PA of the cleanup hitter.
Tuesday, August 30, 2011
A Completely Unnecessary Pitching Metric
There are a number of methods available to evaluate pitcher’s starts on a game-by-game basis rather than the more traditional full season method. There are Game Scores, Support Neutral records, Win Values, and a number of other approaches. There really is no need to add another approach to the mix, and I’m not really going to here--I'm simply going to take a conventional approach for evaluating a full season pitching line and apply it to individual starts.
Of course, I’m not going to claim that this approach is better than the others, because it’s not. It is relatively easy for me to implement, though, and I thought it would be nice to be able to offer a category on my year end stat report for starting pitchers that would consider distribution of performance rather than just aggregate performance as is the case for the rest of the metrics.
The idea is basically to estimate the winning percentage that a team should have over the long haul given the runs allowed and innings pitched of the starting pitcher. I am not using any sort of component RA estimate, and will not bother to explain the implications of this, which I’ll assume you’re well aware of (and thus can also decide for yourself whether you still have any interest in the results). To do this, I use Pythagenpat and assume that the performance of everyone other than the starting pitcher is league average. That is, the offense scores an average number of runs, and the bullpen allows an average number of runs. For the latter, I’m not going to account for the difference between the RA allowed by relievers and the overall league average.
That is not at all an inevitable choice, and if I was trying to construct a perfect metric I wouldn’t do it. But this is so obviously not a perfect metric that the extra effort would be of questionable value. Making that adjustment would also highlight the fact that this approach does not attempt to account for the effect of the number of innings the starter is able to log on the subsequent performance of the bullpen, despite research that suggests there is such an effect. Of course, using the league average doesn’t do anything to address that issue, but it also keeps things simple.
An additional benefit of not making any adjustment for the lower RA of the bullpen is that it allows this metric to be more easily comparable to other metrics that compare the performance of starting pitchers directly to the overall league average--which is a sizeable number. The overall expected winning percentage for the team of a league average starting pitcher at the end of this road will be sub-.500, which while obviously false does in fact match the results of many full season-type metrics.
One thing that cannot be ignored is park effects; the question is how to apply them. One option is to only apply them to the elements of the team other than the pitcher--the bullpen and the offense. I’ll call that option A; option B is to apply the park adjustment only to the pitcher himself.
Option A is a little harder to implement, since there are two adjustments that need to be made. On the other hand, it has some appeal because it allows us to keep the actual run environment of the game rather than recasting it in an imaginary neutral park. I’ve decided to go with Option B because simplicity is a guiding principle here, and because it is more consistent with the way I apply park adjustments to full season metrics. Again, it’s far from an inevitable choice. I’m also assuming that all games are nine innings.
With the thought process out of the way, this isn’t a particularly hard metric to demonstrate. I’ll start simple, with a pitcher throwing a complete game in a neutral park in which he allows zero runs. His expected winning percentage for the game is 1.000.
Seriously, let’s consider a pitcher in a neutral park working seven innings and allowing two runs. I’ll assume it’s an AL pitcher, so we need to know that the 2010 AL average R/G was 4.45 (this is the constant N later). The pitcher’s team can thus be expected to allow 2 + (9 - 7)*4.45/9 = 2.99 runs and score 4.45 runs. This is a 2.99 + 4.45 = 7.44 RPG environment, which has a Pythagenpat exponent of 7.44^.29 = 1.79, and thus the pitcher’s team has an expected W% of 4.45^1.79/(4.45^1.79 + 2.99^1.79) = .671.
We could go through and count up the wins (.671) and the losses (1 - .671 = .329), but I’d rather keep it in rate terms, so the final result will just be the average expected winning percentage across a pitcher’s starts.
To generalize the formula, let N be the league average R/G with R and IP as the runs allowed and innings pitched for the starting pitcher in a particular. Let dPF be the Park Factor without any adjustment so that it can be applied to full season statistics combined for home and road games. For example, the park factors I publish (which are the ones I’ll use here naturally) adjust for this. A 1.03 PF does not mean that the park inflates scoring by 3%--it means that the park inflates scoring by 6%, and is diluted by averaging with 1.00 (neutral park) so that it can be applied to full seasons statistics which are, at least in theory, comprised of one-half home games and one-half road games.
Then:
A (team RA for game) = R/dPF + (9 - IP)*N/9
X (Pythagenpat exponent) = (A + N)^.29
gW% = N^X/(N^X + A^X)
There’s really not much to it when you write it in math rather than English.
I’m very tempted to cap it off by unscrambling it from a W% back into an estimated run average, but I’d rather not deal with the implications of aggregating multiple Pythagorean exponents. One of the advantages of a game-by-game approach is that you’re able to better match performance with the run environment in which it actually occurred, and thus avoid some of the distortions that are inevitable when performance is aggregated across different run environments.
I intend to implement this fully for 2011 starting pitchers (although I might change my mind on that depending on how I feel about the effort/usefulness tradeoff in October), but for now I ran the top five AL starting pitchers from 2010 (IMHO) through the process. For comparison, I’ve included a column called sW% which is based on a traditional use of a pitcher’s full season line to estimate the theoretical W% of his team (albeit without making any adjustment for innings/start):
X = (RA/PF + N)^.29
sW% = N^X/(N^X + (RA/PF)^X)

If you use R and IP as the criteria, Felix Hernandez turned in the best performance of any AL starter, whether you aggregate or consider each game separately. Sabathia and Weaver come out about the same either way, but Lee and Price move in opposite directions when you look at the game level. This implies that Lee’s distribution of runs allowed and innings was such that it would figure to produce more wins than the averages would suggest, with Price the opposite.
This is more of a freak show stat than anything else, but it does provide a relatively simple way to compare starters at the game level on their bottom line results, and if a Cy Young race is particularly close, you may want to consider it. Or you may not; I don’t have a lot of conviction about this, and there are more rigorous approaches available, but there you go.