Showing posts with label Book Reviews. Show all posts
Showing posts with label Book Reviews. Show all posts

Saturday, March 06, 2021

dWhat!%

It’s understandable that the editing process for Baseball Prospectus 2021 overlooked something trivial like explaining what a metric in the team prospectus box means. After all, it must have been exhausting work to ensure that each of the many political non-sequiturs in the book were on message (Status: success! You can give this book to your children to read with confidence that they are in a safe space, with no deviation from the blessed orthodoxy). The vital imperative of ideological conformity handled, they would have needed next to run a fine-tooth comb over any reference to the aesthetics of present day MLB on-field play to ensure the proper level of smug conflation of one’s own preferences with the perfect ideal. Another success. Finally, they could turn their attention to making sure there were the requisite number of sneering statements about the fact that there even was a MLB season in 2020.  As always, left unaddressed was how a publication that exists (in theory at least – reading the 2021 annual, this may be a fatally flawed assumption on my part) to analyze professional baseball could continue to exist if professional baseball ceased to exist, but who knows? When you tow the line so perfectly, maybe you can figure out a way to get in some of that sweet $1.9 trillion.

So it is entirely understandable that such a triviality as a publication rooted in statistical analysis could completely overlook explaining a metric that none of its writers ever bother to refer to anyway. The metric in question is called “dWin%”. It didn’t replace any team metric that was listed in the 2020 edition – it literally fills in a blank space in the right data column. A search of the term “dWin%” and “Deserved Winning Percentage” on the BP website doesn’t yield any obvious (non-paywalled, at least) relevant hits. So the best I can do is make an educated guess about what this metric is.

I gave away my guess by searching for “Deserved Winning Percentage”. BP has adopted a family of metrics with the “Deserved” prefix which utilize Jonathan Judge’s mixed model methodology to adjust for all manner of effects (going well beyond the staples of traditional sabermetrics like league run environment and park). The team prospectus box lists “DRC+” and “DRA-“, which are the DRC metric for hitters and DRA for pitchers indexed to the league average. So it’s only natural to assume that dWin% is some type of combination of these two to yield a team’s “deserved” winning percentage.

It’s also natural to assume that there would be a relationship between DRC+, DRA-, and dWin%. If the first two are in essence run ratios (with myriad adjustments, of course, but essentially an estimate of percentage difference between a team’s deserved rate of runs scored or allowed and the league average), then it’s only natural to assume that there would be some close relationship between them and dWin%. If we were in the realm of actual runs scored and allowed, or runs created/runs created allowed, we could confidently state that one powerful way to state the relationship would be a Pythagorean approach. Namely, the square of the ratio of DRC+ to DRA- should be close to the ratio of dWin% to its complement.

There are two obvious caveats to throw on this conclusion:

1) While the statistical introduction does not specifically refer to DRA- (it refers just to DRA, which was listed for teams rather than DRA- in the 2020 edition), it’s reasonable to assume that DRA- is the indexed version of DRA. DRA is a pitching metric, which would attempt to state a pitcher’s deserved runs allowed after removing the impact of the defense that supports him. This means that comparing the ratio of DRC+ and DRA- on the team level is likely ignoring fielding, and thus the relationship I’ve posited above would be incomplete. I would be remiss in saying that this is not the fault of BP, except to the extent that we are left to speculate about the meaning of these metrics, as there's certainly nothing wrong with having a measure that attempts to isolate the performance of a team's pitching staff.

2) It is possible that there is something else going on besides fielding in the process of developing the Deserved family of metrics that would invalidate this manner of combining the offensive and pitching components. Without being privy to the full nature of the adjustments made in these metrics, it’s hard to speculate on what if anything that might be, but I would be remiss in not raising the possibility that there’s something going on behind the curtain or that I have simply overlooked.

I’m not going to run a chart of all of the team values, because that would be infringing on BP’s property rights, and given the first paragraph of this post that would be practically unwise even if it were not morally objectionable. A few summary points provide defensible ground:

1) the average of the team DRC+s listed in the annual is 99.3 and the average of DRA-s is 99.5. Given that the figures are rounded to the nearest whole number (e.g. 99 = 99%), this is encouraging as we would expect the league average to be 100.

2) the average of the team dWin%s is .464. Less encouraging. As I was reading through the book, there were two team figures that really caught my eye and led me to this more formal examination. The first was Philadelphia, which had a dWin% of .580, ranking second in MLB. Their DRA- was 83, also second.

The Deserved family of metrics have always produced some eyebrow-raising results, which are difficult to evaluate objectively given the somewhat black box nature of the metrics and the complexity of the mathematical approach involved (I will be the first to admit that “mixed models” of the kind described are beyond my own mathematical toolkit). So it’s dangerous to focus too much on any particular result, as it may just be a vehicle by which to expose one’s own ignorance. As a second-generation sabermetrician, this is a particular nightmare, becoming the sportswriter you laughed at as a twelve-year old for dismissing RC/27 as impossibly complex and unintelligible.

Still, it is quite remarkable that the team which allowed the second-most park-adjusted runs per inning in the majors might actually have turned in the second-best performance. In fairness, it was a sixty-game season, so the deviation between underlying quality of performance and actual outcome could be enormous, and the East could have been the toughest of the three sub-leagues, especially in terms of balance as the Dodgers tip the scales West. Most significantly, it is just a pitching metric, and the Phillies defense was dreadful at turning balls in play into outs – they were last in the majors in DER at .619. Boston was at .623 and the next worst team was Washington at .642. Further, the East subleague combined for a .657 DER (the fourth-worst DER belonged to the Mets, and Toronto and Miami made it six of the bottom ten) compared to .685 for the Central and .684 for the West. It’s still hard to believe that the Phillies’ pitchers deserved to have the second-fewest runs allowed in the majors, but easy to buy that they performed much, much better than their runs allowed would suggest.

However, every factor that would explain how their pitching was actually second-best does nothing to explain how their overall deserved team performance was also second-best. Adjusting away terrible defensive support doesn’t mean that the team’s poor runs allowed weren’t deserved, it just means that the blame should be pinned on the fielders and not the pitchers. Again, it’s hard to pinpoint any exact criticism given the nature of the metrics, but this one is tough to accept at face value.

It also seems that if one had conviction in the result, it would show up in the narrative somewhere. There’s always been a disconnect between what BP statistics say and what their authors write, which owes partly to the ensemble approach to writing and presumably partly to the timing (the authors of team chapters probably start very soon after the season and without the benefit of the full spread of data that will appear in the book). Still, it seems as if this disconnect has increased with the advent of the deserved metrics, which often tell a very different story than even the mainstream traditional sabermetric tools (e.g. an EqA or a FIP, to refer to metrics previously embraced by BP). But I can assure you that if I believed the Phillies underlying performance as a team was actually second only to the Dodgers, I’d work that into any retrospective of their 2020 performance and forecast of their 2021.

The second team that caught my eye was the A’s, who posted a 103 DRC+, 98 DRA-, and .499 dWin%. The obvious disconnect between an above-average offense, above-average pitching, but sub-.500 deserved W% could be explained by defense. What can’t be explained is how a .499 dWin% ranks ninth in the majors, at least until you line up the thirty teams and see that the average is .464. While we can charitably assume that a combination of our own ignorance and the proprietary nature of the calculations can explain many odd results from the deserved stats, I don’t know what can satisfactorily explain a W% metric that averages to .464 for the whole league.

The hope is that this simply some scalar error, a fudge factor not applied somewhere. There is some evidence that this is the case – if you take the ratio of DRC+ to DRC- and plot against the ratio of dWin% to (1 – dWin%), you get a correlation of +0.974 and a pretty straight line, as you would expect given what should be in the vicinity of a Pythagorean relationship. It might even work out as you’d expect if dWin% is baking in fielding.

Still, it’s disappointing that the question has to be asked.

Wednesday, May 18, 2016

The Only Rule Is It Has to Work

Note: The following is a rare (for this blog) timely book review.

The premise of The Only Rule Is It Has to Work is that respected sabermetrically-inclined authors and podcasters Ben Lindbergh (Baseball Prospectus, Grantland, Five Thirty-Eight) and Sam Miller (Orange County Register, Baseball Prospectus) were given the opportunity to act as the baseball operations department throughout the 2015 campaign of the Sonoma Stompers, member of the four-team Pacific Association, a low-level indy circuit in northern California. Lindbergh and Miller were granted wide berth to put their mark on player acquisition, roster construction, and in-game strategy, and also attempt to bring modern data collection tools (PITCHf/x, video scouting, etc.) to the bush leagues.

Lindbergh and Miller are embedded deep within the team--in the front office, the clubhouse, the dugout, and even (for a moment at least) kangaroo court. Thus it serves as one of the most revealing examinations of daily life in baseball from an outsiders' perspective. Most books that have provided similar access to the inner workings of a team have been written by insiders, even if they might not fully fit into the world in which they have spent many years (think your Jim Boutons). While the life of an indy-league player is certainly less lavish than that of a big leaguer and perhaps less structured than that of an affiliated minor leaguer, it's hard to imagine that the basic human impulses of (largely) twentysomething, athletically-gifted ballplayers varies much between Sonoma and San Jose, San Jose and San Francisco. The authors are able to observe the scene with some combination of bemusement, paternal-ish concern, and comradery to give the audience a different perspective on the people who play the game. Certainly the majority of the audience members can better relate to the authors' stations in life and can now imagine how they might fit in (or not) if thrust into the life of a ballclub.

While it should hardly be necessary at this point for sabermetricians to defend themselves against scurrilous charges of not watching the games, one thing that the authors don’t reflect too closely upon but that is obvious to the reader is just how much low-level baseball they watch over the course of the summer, and just how devoted to their cause they are. Granted, Lindbergh and Miller are aided by a small network of volunteer scouts that earns the derisive nickname "The Corduroy Crew" around the league, but one or the other personally does advance scouting of nearly every game the Stompers' opponents play. This in addition to the hours spent researching potential players with their proverbial noses buried in a spreadsheet. While it would be wrong to hold Lindbergh and Miller's labors (which of course were performed with at least the secondary intent of providing fodder for a book) up as a pure representation of baseball love to be extrapolated to all of their sabermetric compatriots, it would be less wrong to do so than to brandish the common stereotype.

One of the disappointments of the project is that many of the radical ideas the authors dreamed about being able to test are never put into play. While shifts and flexible usage of the relief ace take hold in the second half of Sonoma's season, batting orders largely remain tethered to convention, starting pitchers still generally work in rotation, and the manager holds on to ultimate in-game command. While this may be disappointing to the reader longing for sabermetric red meat, the implications raise questions worth considering. Is it necessary for change in baseball tactics to come one easily digestible piece at a time? Why can a grizzled bench veteran and former pennant-winning manager of a major league team (Clint Hurdle) pivot to the approach his superiors' desire with more aplomb than a 37-year old pot-smoking player-manager who goes by Feh and dabbles in 9/11 conspiracy theories? Do the high stakes of the majors actually make them a more suitable laboratory for experimentation, as players and managers can count on their million dollar checks regardless of whether they may look unconventional on the field? While these questions can't be answered by the book, it provides some entertaining anecdotal evidence to consider.

Along the way, the Stompers inadvertently break ground in the social realm of baseball as well, as one of the authors' hand-picked college signees, relief ace Sean Conroy, comes out as the first openly gay player in professional baseball. The authors do an excellent job of relating this part of the story without falling into self-congratulations or allowing it to swamp the baseball portion of the narrative. Lesser authors with a less interesting baseball story to tell (and perhaps less respect for their subject) could have easily allowed Conroy's story (which includes being one of the Pacific Association's most valuable pitchers) to crowd out other aspects of the Stompers' season in the narrative, and could hardly have been blamed for it.

The authors alternate chapters, and if you are a regular listener (as I am) of their Effectively Wild podcast, you will likely be able to pick out which voice you are reading after a couple of pages even if you forget for a moment whether it is an odd or even chapter. Lindbergh's earnest verbosity and Miller's cheerful nihilism carry through to the written page in book format yet complement each other well, imbuing a diversity of style to the writing while still making you feel as if you are reading the same book.

As luck (or the residue of design) might have it, the story has a dramatic conclusion that I will not spoil here, except to say that I'm very glad the majors have resisted the allure of the half-season format, except for every ninety years when unusual circumstances take hold (if I live to see baseball in 2071 I promise to be grateful and not complain about it too much). Were it ever turned into a movie, the scriptwriter would even have something of a "pick your own adventure" opportunity to affect the outcome with only the proverbial flap of a butterfly's wing.

And maybe that's one of the lasting lessons to take away from The Only Rule Is It Has to Work. That despite the careful planning, the on-the-fly adjustments due to injuries or player poaching (at this level), the dedication of the players and support staff, the superstitious rituals, and the motivational speeches that are poured into baseball clubs, not to mention the attempts to drag baseball kicking and screaming into the sabermetric age, we will never be able to escape what seem from our imperfect perspective to be random rolls of the die.

Monday, May 04, 2015

Counter-Revolution

Continuing the tradition of haphazard “book reviews” appearing on this blog well past the time that such a review would be relevant, I recently read The Sabermetric Revolution by Benjamin Baumer and Andrew Zimbalist and have a few thoughts on the book.

On the whole, I am not a fan of the book. While I am not personally very familiar with Baumer’s work, Zimbalist is a seminal figure baseball economics (starting over twenty years ago with his Baseball and Billions). Unfortunately, The Sabermetric Revolution is too short (153 pages of prose not counting footnotes) and too unfocused to really showcase the authors’ knowledge.

In many respects it appears that the book was intended to be something of a rejoinder to Moneyball, both by pointing out areas in which Michael Lewis either played fast and loose with the facts or omitted key details. The preface is clear about this motivation, as the authors write: “This book will attempt to set the record straight on Moneyball and the role of ‘analytics’ in baseball.”

There’s no doubt from reading the book that this is a major goal of the authors, as the first chapter is devoted to “Revisiting Moneyball”. I found some of the criticism to be fair (for example, Lewis’ tendency to gloss over the contribution of young talent the A’s had produced that contributed to the team’s success, such as Eric Chavez, Miguel Tejada, Tim Hudson, Mark Mulder, and Barry Zito). Some, though, strikes me as of 20/20 hindsight (such as a review of the infamous 2002 amateur draft) or nit-picking (such as the fact the A’s OBA decreased in 2002 despite Beane’s emphasis on OBA). In other places, I would contend the authors are guilty of some of the same offenses they accuse Lewis of (for example, they state that Lewis gives short shrift to the work of Bill James and other sabermetric pioneers; however, their own discussion of the internet sabermetric community begins at Baseball Prospectus).

The fundamental issue I had with the book is that it is not clear what it is intended to be (aside from a Moneyball response) or who the intended audience is. The book is not detailed enough to serve as a technical introduction to sabermetrics for newcomers (for instance, I’m not sure park factors are ever discussed outside of brief allusions), but neither is it detailed or advanced enough to strongly appeal to the smaller audience of practicing sabermetricians. There is even a chapter on statistical analysis in other sports, a topic on which I am closer to the novice group, but it also is short on details, even more glaring of an omission since at least there is a quick overview of sabermetric theory.

At the cost of myself falling into the trap of nit-picking this book, I think listing a number of my issues with the book might be the easiest way to write it up:

* There are also a number of incorrect acronyms used in the book, some of which were surprising to me. OPS is said to be an acronym for “Offensive Performance Statistic”; DER an acronym for “Defensive Efficiency Rating”.

* The authors state that the formula for Isolated Power weights doubles and triples equally and is roughly the difference between SLG and BA, “or sometimes” is (D + 2T + 3HR)/AB. While I understand the argument for treating doubles and triples equally in a power metric, Isolated Power is not “sometimes” defined as SLG minus BA. That formulation has been used in conjunction with the term “Isolated Power” since Branch Rickey linked the two (but did not set them equal) in his 1954 Life magazine article and it was used in the manner by Bill James. While this or the meaning of the OPS acronym may seem like insignificant details, they suggest something less than a full command of sabermetric history.

* The authors state that in economic terms, WAR measures “marginal physical product” and state that this is a good idea, but are not fans of the methodology used to calculate current WAR implementations. Their concerns include fair ones, such as failure to report error bars and the use of black box methodologies. But while their reasoning behind these criticisms are clearly laid out, they sometimes engage in what might be called “drive-by” criticisms, in which issues are alluded to but not fully fleshed out to the point where the creators and users of these metrics could offer a defense. In this manner, Baumer and Zimbalist reflect the attitude of another “insider” who has criticized replacement-level metric, Christopher Long.

One such comment is “It is not clear that there exists a pool of replacement players with the productivity that is ascribed to them”. This basically questions the entire concept of replacement level, but is not supported other than with a footnote to site the work of JC Bradbury. This does nothing to forward the discussion of replacement-level, nor does it alert the readers to the well-reasoned and spirited rejoinders sabermetricians have issued to Bradbury’s contentions.

The authors then use a single example to question what is one of the least controversial and most similar step in any WAR methodology--the run to win conversion. The authors simply write: “The use of James’ Pythagorean Expectation to convert runs to wins is less than robust. One need only reflect on the 2012 Baltimore Orioles, who outperformed their expected win total by 11 games, to see how inaccurate the runs to wins conversion can be.”

If I may be impolitic and a bit unhinged for a moment, the authors should be ashamed of themselves for this statement. It is the type of statistically illiterate cherry-picking that one might expect from a Bill Madden rather than from respected professionals familiar with statistical methods. While it is without question true that win estimators (like every other statistical estimator known to man) produce poor estimates in certain individual cases, a reasoned discussion of their error bars does not begin and end with a single poor estimate. Any regression equation presented by Zimbalist in Baseball and Billions or in this work could be easily impugned by similar rhetoric, and likely more effectively given that win estimators are among the more accurate and stable estimates one will find in baseball analytics.

It might also be pointed out that the run to win converters actually used in WAR calculations are likely more robust (in the true meaning of the term, rather than denoting a single outlier) than Pythagorean by recognizing that the shape of the relationship between runs and wins changes as the scoring environment changes. While the authors are surely aware of this, one could never tell from the discussion of run/win estimators in the book, as only Pythagorean constructs with fixed exponents are discussed, with no reference to alternative exponent constructions like Pythagenport/pat or dynamic linear run to win estimators.

* My sense, and it may be unfair, from reading the book, is that Baumer and Zimbalist are eager to emphasis areas and issues in which sabermetric findings have been wrong and/or incomplete. An example is the discussion on sacrifice bunts, which points out that the initial sabermetric analysis (they do not reference Palmer and Thorn by name in this section, but The Hidden Game of Baseball is the usual source of the classical argument) was incomplete in not considering the other outcomes that may occur on a sacrifice bunt attempts, such as bunt hits and errors.

This is without question a valid criticism. However, neither Baumer/Zimbalist nor other present day critics of the conclusion acknowledge that the conventional wisdom that was pushed back against was not that the bunt was a good play because of those outcomes, but that the sacrifice if successfully executed was a good play. I still find myself as one of the few patrons clapping when I attend a game and the team for which I am rooting successfully records the out at first base on a sacrifice. This play was seen, and still is seen by casual fans and presumably a non-negligible portion of major league managers, as a success for the offense, even without the benefit of the error or hit that make the play a palatable strategy in certain situations. Sabermetricians have moved to a more “nuanced understanding” of the sacrifice, but they have also forced the conventional wisdom to tack on a bunch of addendums and hypotheticals that had rarely been discussed before.

* In other cases it is unclear how deep of a literature review of the field the authors have performed. For instance, the authors criticize FIP due to using an ERA scale (a criticism with which I agree but also note can be relatively easily corrected) but state that “What this field needs is a simple, illustrative, but effective model to evaluate pitchers. Until a model can be constructed with interpretable coefficients (a la linear weights), or with meaningful interaction of terms (a la Runs Created), no real insight will be gained, and there is unlikely to be any consensus about which metric is best.”

In all,The Sabermetric Revolution is a book that I think might have been better conceived as a couple of separate journal articles on the topics on which Baumer and Zimbalist have something new to say, because the rest of the book feels like filler and does not establish a consistent purpose or tone.

Tuesday, June 14, 2011

Comments on Bill James' “Solid Fool’s Gold”

For the past three seasons, Bill James had published an annual book called the Bill James Gold Mine. The book included a sampling of some of the material available to subscribers of his Bill James Online website, including some unique split data (unique in the sense that it’s not commonly in print on actual pieces of paper), like breakdowns of pitches thrown to left-handed and right-handed batters. There were little boxes with “nuggets” (gold is a theme needless to say) pointing out various oddities.

Those elements of the book really took up a lot of space, but were woefully incomplete (they clearly weren’t intended to be complete, but the point is that the book had no utility as a reference). The most interesting aspect of the book was that it reprinted a number of full-length essays that James had written for his website over the course of the year. I generally enjoyed those essays as you will see if you look back at my comments on the previous editions of the book.

This year, there was no Gold Mine. Instead, ACTA and James released a slimmer, smaller volume entitled Solid Fool’s Gold, which includes only essays. I didn’t bother to count, but I’d guess that the new book has about as many essays as the Gold Mine, it is sold for a lower price, and personally I won’t miss the hodgepodge of charts that much.

If a fairly quick read of non-technical essays by Bill James is something you think you’d enjoy, you’ll probably like Solid Fool’s Gold, unless you already subscribe to Bill James Online. The new format is much better, as it includes a bunch of essays that you could pick up and read in five years rather than some of the more limited shelf-life aspects of the Gold Mine and it does it more cheaply and compactly. I certainly enjoyed it, which you should keep in mind as I now launch into a more critical review of some specific essays in the book.

The essay that has probably gotten the most attention (outside of the much-panned essay on Shakespeare and Topeka which was published on Slate) is called “Minor League Pyramid”, and it includes James’ outline of a way to reform the minor league structure to make it resemble a pyramid rather than a tube (his description) as it does now. I won’t repeat his argument here, but there are couple points that I feel strongly about:

1. I agree with James that talent is choked out of the game by the limited number of openings at the entry level. This is offset somewhat by the existence of college baseball as an alternative means to improve baseball, but scholarship restrictions make it a less reliable means of keeping quality talent engaged than college football or basketball. The scholarship restrictions also leave college baseball as next to useless for attracting low-income players to the game.

2. The proposal that James offers includes limits on the rate at which a prospect can be advanced through the minors, with the bottom line being that a player could not reach the majors until he’d played three seasons in the minors. This is impossible to square with the existence of college baseball, and it also removes the illusion of a meritocracy which I think is very important, even if it is only an illusion. We all know that teams play service time games, and that there are few twenty-year olds ready to be major league contributors anyway, but the potential for a wunderkind to reach the majors, even if there are only a handful each year, is something that should be preserved.

3. James doesn’t seem to think his system would have much of an impact on the ability of baseball to attract talented athletes with other options. In fact, he claims that a minor league pyramid would reduce the pressure on signing bonuses. While I’m sure MLB CFOs would like that (as they would like the destruction of college as an alternative path), I can’t fathom how it wouldn’t make MLB a much less attractive option.

He doubles down on this by saying that it would reduce “pressure” on teams to scout internationally to find talent, since they would have a larger supply of homegrown players. I can’t for the life of me understand why that would be considered a positive. The piece does make some good points, but that one is a real head-scratcher.

Another essay in the book is called “Stink-O-Meter”; it discusses a fairly simple method of tracking the persistency of losing for a franchise. The article is good in that it reinforces something that the average fan with no historical perspective constantly needs to be reminded--the state of even sorry franchises like the Pirates and Royals is nothing like that of the terrible franchises of the game’s past.

The article is way too long, though--James feels compelled to run a chart every few paragraphs listing the top five or ten losing teams at a given moment in time. This works better online than in print where it just wastes space, but it also helped me to see what I think is a pattern in James more recent work. James has realized that it’s more difficult (not just for him, but for anyone) to produce cutting edge technical work. Rather than introducing more rigorous means of analysis, James has decided to play show-and-tell; his more recent essays are filled with tables that in the past would have been left to the imagination or determination of the reader. I could be off base, but it seems that increased comprehensiveness represents James’ attempt to keep up with the Joneses.

Another essay is a reprinting of a speech James gave called “Battling Expertise with the Power of Ignorance”. It’s a good read, but there was a portion that mentioned Pythagorean record and Runs Created that, shockingly, I can’t help but comment on.

James describes the two methods as the best-known of the “large number of heuristic rules” he developed during what could loosely be defined as his Abstract years. About the Pythagorean theorem, James said: “Later research has demonstrated that it works better still if you modify the exponent for the level of scoring”.

The relationship between the slope of a run to win converter has been known for many years (dating back at least to Pete Palmer), but Clay Davenport’s Pythagenport was the first well-known modification to James’ Pythagorean theorem. Later, this was refined further into Pythagenpat by recognizing the minimum theoretical exponent was one.

James freely acknowledges the refinements to Pythagorean and in fact has used Pythagenpat in at least one of his own studies. That only makes it all the more strange that he continues to cling to Runs Created, even when RC has been demonstrated to be a less accurate tool than Pythagorean with a fixed exponent. RC is subject to complete meltdown under theoretical extreme conditions; while Pythagorean incorrectly handles the known point of 1 RPG, it correctly imposes a range of [0, 1] on all of its estimates.

Discussing RC, James made no mention of subsequent work on other run estimators. Understand that I am not trying to claim that he had any obligation to do so in the context of this speech--only that his continuing clinging to RC while recognizing other refinements to his original tools grows more bizarre as time passes. I suppose one could argue that at least the Pythagorean refinements maintain the original R^x/(R^x + RA^x) model, and James always pointed out that an exponent other than two could result in more accurate estimates. In any event, it’s extremely hard for me not to comment on a run estimator when the opportunity arises.

While James recognized the existence of variable exponent refinements to Pythagorean record, he unfortunately missed a golden opportunity to utilize them in another essay in the book. There is an article in which James examines the performance of starting pitchers when supported by X runs--basically, an attempt to examine the mystical phenomenon of “pitching to the score”. I won’t steal his thunder by discussing his conclusions, but I will point out a methodological shortcoming in his approach.

When Whitey Ford’s teams scored one run with him pitching, their record (not Ford’s record) was 10-28. James converts this to an effective rate of runs allowed using Pythagorean math. In this case we know that Ford’s teams scored 38 runs (one for each game), so the equivalent number of runs Ford allowed to produce a Pythagorean record of 10-28 is x in the following equation:

10/(10 + 28) = 38^2/(38^2 + x^2)

This eventually simplifies to sqrt(L/W)*R, or sqrt(28/10)*38 = 63.6. With two runs, Ford’s team were 19-22, which is equivalent to sqrt(22/19)*82 = 88.2 runs. Adding these up and dividing by the total number of games produces an “Effective Runs Allowed Rate” for Ford in games in which his team scored one or two runs: (62.6 + 88.2)/(38 + 41) = 1.91. Continuing in this manner for scoring three, four, five, … runs (while ignoring shutouts which are always losses), James has a measure of pitching effectiveness given the level of offensive support on a discrete game-by-game basis.

However, the use of a fixed exponent severely distorts things by essentially assuming an average run scoring environment (an exponent of 2 corresponds to a RPG of around 10.9 using Pythagenpat), when we know the scoring output of one of the teams involved. If one team scored only one run, the expected RPG is going to be lower than average.
We could assume that an average number of runs would be scored by the other team involved in the game, and instead say that the RPG is 5.5, and the Pythagorean exponent should be around 1.64. In that case, the equivalent runs allowed would be (28/10)^(1/1.64)*38 = 71.2 runs, a 12% difference from James’ estimate.

That approach assumes that we know nothing about the “other” team’s run scoring rate--but of course, we know a great deal about it, because we know the identity of the starting pitcher: Whitey Ford. For his career, Whitey Ford had a Run Average of 3.14 and averaged 6.94 innings/start in a league that averaged about 4.31 runs/game, so we could estimate that his team’s RA is (6.94*3.14 + (9 - 6.94)*4.31)/9 = 3.41, and that the expected RPG for a game in which his offense scores one run is 4.41, producing a Pythagorean exponent of 1.54 and (28/10)^(1/1.54)*38 = 74.2 run equivalent. This new estimate is approximately 17% higher than James’ original estimate.

The good news is that when you extend things across the entire spectrum of the run distribution, much of the distortion is canceled out. James presents complete breakouts for several pitchers, but the two of historical interest are Ford and Tom Seaver. Setting aside the adjustments he introduces to smooth the data and restate the effective runs allowed rate on the actual RA scale, let me just run the crude totals for those two under three assumptions: the James approach of a Pythagorean exponent of 2, an assumption that the RPG at each scoring level is 4.5 plus the number of runs the pitcher’s team scored, and the customized type assumption I described for Ford above (Seaver had a career RA of 3.15 and averaged 7.38 innings/start in a league that averaged approximately 4.11 runs per game, resulting in a customized team RA of 3.32 runs/game):



These differences are small, in the neighborhood of 1%, and thus not worth getting too worked up about. However, it’s important to keep in mind that fixed Pythagorean will not work particularly well at the extremes, and it would be a mistake to put a lot of confidence in the isolated application of the fixed exponent Pythagorean estimate to an extreme RPG.

Thursday, May 19, 2011

Meanderings

Print Baseball Encyclopedias

As I grow older, I try to stay alert to warning signs of old-fogeyism. One or two such signs are not particularly concerning--they can just be written off as personal quirks/eccentricities, which we all possess to one degree or another. A prime example for me is cell phones. I hate the things, and I always have. I finally got one, only because it was cheaper than paying for a landline, and if there's one thing I hate more than cell phones, it's spending money on any type of phone.

When it comes to baseball, one of the possible signs I've noticed is my continuing love for print encyclopedias. I think it's great that we have Baseball-Reference, Retrosheet, the National Pastime Almanac, the Baseball-Databank, and the like, and obviously there are countless advantages to computerized data that you and I take advantage of every day. Still, I have yet to warm up to the idea of going to Baseball-Reference, clicking on a page, following a link somewhere else, and wasting an hour or two just wandering in the statistical record of the game. I still do this all the time with print encyclopedias. This post is a tribute/review of them.

Of course, the print encyclopedia is a dinosaur. It always was a bit of a wonder that one could publish a multi-thousand page book, carrying a hefty hardcover price, and sell enough of them to make it a worthwhile business endeavor, especially with annual or semi-annual editions. Perhaps they never really earned their keep anyway, but they should have.

The advent of computerized equivalents has driven the print encyclopedia out of existence (although apparently the erstwhile ESPN Baseball Encyclopedia is still being shopped to publishers). If that is the inevitable cost of progress, then so be it--I wouldn't give up my Lahman database to get a new edition of Total Baseball if that was what it would take. Still, I miss the print encyclopedias--and it seems as if other people do to.

As I write this (New Year's Eve), the current cheapest prices listed on Amazon.com for a copy of the final edition of each of the printed encyclopedias (new or used) are:

* Macmillan (10th edition, 1996): $44.99

The 9th edition is available for as little as $25.

*Sports Encyclopedia: Baseball (2007 edition): $123.08

The 2006 edition is available for as little as $3.31.

*Total Baseball (8th edition, 2004): $99.65

The 7th edition is available for as little as $3.58.

*ESPN Baseball Encyclopedia (5th edition, 2008): $95.80

The 4th edition is available for as little as $1.73.

*STATS All-Time Baseball Handbook (2nd edition, 2000): $3.99

The exception, and not really an iconic book as it only went through two editions and presumably had the most limited printing run of any of the five.

I'm not sure if these prices reflect actual demand for the books in question, or whether sellers think they have something valuable and are setting the price above the intersection of the demand and supply curves. Assuming that it is a real phenomenon, it suggests that there are a fair number of people who miss the print encyclopedias so much that they are willing to pay a high price just to have the final update.

I have at least one copy of each of the big four (excluding the STATS book from that designation) on my bookshelf at all times. Of the four, the two that I use most are ESPN and Sports Encyclopedia: Baseball. Of all of the encyclopedias, I have to count SE:BB as my favorite. It's certainly not the most statistically complete or the best-edited, but it's the only one of the four that breaks from the career register format and instead presents a season rosters format.

I've always felt that the season rosters lend themselves better to browsing than the career registers. (This is the part where the readers scream, "With a computer you can have both!") Not only does it allow one to look at team composition and track changes from year-to-year, it allows one to view an entire league-season on 2-4 pages, making it much easier to get the big picture for a season.

The SE:BB is not without flaws, of course. The book is filled with typos, many of which were presumably there from the first edition to the last. Two quick examples, both from the 1994 edition (although I'd be very surprised if they were corrected in later updates):

* Johnny Kling is listed as "Johnny King" with the roster for the 1901 Cubs, and in the 1901-19 Batter Register (later Cub seasons correctly list him as "Kling").

* The header for the 1972 NLCS says "Cincinnati (west) 3 Pittsburg (East) 2". Perhaps if this was a listing for 1882, it could be considered authentic to the times.

There have to be dozens of similar errors throughout the book, none of which are damning to its utility as a baseball reference but all of which do build up to an uneasy feeling of neglect. Still, the charms of the book overcome that for my money.

Like its cousin, the Macmillan, the statistical selection in SE:BB was formed at its first publication (1969 for Big Mac, 1974 for SE:BB). OBA is nowhere to be found, nor is CS or pitcher home runs allowed. Fractional innings pitched are rounded, an the typesetting varies throughout the book, making some sections more difficult to read. Sometimes space requires severe truncating of batting lines--Dick McAuliffe went 7-27 as a 20 year old left-handed hitter for the 1960 Tigers, but that's all you can find out.

The ESPN encyclopedia, edited by Gary Gillette and Pete Palmer, is my favorite of the three career register works. Mostly this is because it is the most recent, superceding Total Baseball. For the most part, the statistical selection is the same as TB. In both cases, I'd love to have a better offensive rate than OPS+, and I think they tried to hard with respect to fielding categories, but both give the basic categories necessary to build standard statistics.

Total Baseball is unique because of the volume of the text that accompanies the statistics--short biographies of notable players, team histories, a history of sabermetrics, and a bunch of other articles that changed from edition-to-edition. More than any of the other encyclopedias, the article turnover created a reason to buy each new edition (other than, of course, the updated statistics).

The MacMillan must be given respect due to its status as the pioneer; the research that went into producing it has been incorporated by every serious baseball historical work of any stripe since that time. As an encyclopedia, though, it's heyday was the first edition. It soon had SE:BB as a competitor, and with the two including essentially the same basic data, the (IMO) superior format of SE:BB made it an unfair fight. MacMillan also played fast and loose with changing statistics for silly ends. Later editions cut this out, and added some interesting data like team home/road splits and sketchy Negro League records, but by that time Total Baseball was on the scene.

The STATS All-Time Major League Handbook was the most thorough encyclopedia for individual statistics, but as such it is the one that has taken the biggest hit from the existence of Baseball-Reference. No other encyclopedia offered complete batting, pitching, and fielding data (including all of the minor categories like GDP and sacrifice hits allowed), but the sheer volume of data sapped the book of any character it might have otherwise had. While Big Mac has standings and playoff records and the like, and Total Baseball had all of that and the articles, there was no room in the Handbook for anything other than the player career register. The ancillary material was shuffled off into an equally large All-Time Sourcebook.

While the massive print encyclopedia may be something of a relic, I do think it would be wonderful if it could live on. Obviously I know nothing about the real-world feasibility of what I am about to spout, but it would be great to see an organization like SABR step up to the plate and subsidize an updated print encyclopedia (even if it had to be in PDF format, as SABR has done with the Emerald Guide) every half-decade or so. Eventually the desire for such a tome might be foreign to even the crustiest old baseball historians, but I think it's safe to say that day is still several decades off into the future.

Standard Deviation of Franchise W%

Speaking of electronic encyclopedias, this is the type of exercise that they make a breeze, which previously would have been an arduous chore. I figured these a while ago with the intent of using them in some other discussion, but that never materialized so I'll dump them here.

These charts simply show the standard deviation of full-decade W% for each major league franchise. I have criticized the use of decades as a line of demarcation for baseball statistics in the past, but this is not a through analytical endeavor and they do provide an easy, straightforward manner of categorization. I have defined the decade here as 1901-1910, 2001-2010, etc, not because I have any particularly strong feelings on the matter of decade division but because it works better since 1) it includes 2010 and 2) the first decade thus defined corresponds with the American League's 1901 ascension to major league status.

There are four different standard deviations shown for each decade--"whole" is the StD for teams that completed the entire decade. This is fairly arbitrary, as it allows the 1961 AL expansion teams but excludes the 1962 NL expansion teams (the four 1969 expansion teams are obviously excluded as well). "All" is the StD for all franchises that played in the decade, even if it was for as little as one season (actually, the shortest in-decade tenure is two years for the 1969 expansion teams). "1901s" is the StD for the sixteen franchises that have played continuously since 1901. While they now make up just over half of MLB, they at least provide a constant frame of reference throughout the century. "Expan", as you might figure, is the StD for whichever of the fourteen expansion franchises competed in a given decade.



By this measure, the 1980s and 90s stand out as very competitive periods in the game, and the 2000s were a step back from that. However, the standard deviation of franchise W% in the last decade were essentially the same as the 1950s and 60s, and still well under the norm for most of history.

The next chart gives the average W% for teams by decade broken down into 1901s and expansion teams. It also lists the best and worst franchise W%s for the decade, but those lists include only the teams that played ten seasons in each decade:



In the 1980s, expansion teams actually had a slightly better record than the 1901s, but they have lost ground in the last twenty years. Of course, most of the big city teams are 1901s, with the major exception being the Angels. The spread between the best team W% and worst was higher in the 2000s than it had been since the 1960s, but I wouldn't attempt to make anything out of it.

Two Team Cities

During a bout of encyclopedia browsing, I noticed that the two Boston teams both had dreadful 1906 seasons. The Braves were 49-102, but the now-Red Sox were even worse, losing three more games (49-105). I made the mistake of pointing this out on Twitter and saying that it "had to be the worst" such record.

Of course, it didn't have to be anything, and it isn't. It is only the third-worst combined record by teams in the same city since 1901. While I'm sure someone has done this before, a quick search turned up nothing. I considered Brooklyn to be New York (meaning that from 1903-1957 New York had three teams), and I considered the Angels/Dodgers and Giants/A's as sharing a city (when applicable). The ten worst single season records for the two or three teams combined:



At least Boston 1906 was the worst in something, as it was the worst non-Philadelphia combined record. Philly has seen some bad records over the years, but none worse than 1919 when the Phillies were 47-90 and the A's were 36-104. The worst years for each of the two-team cities other than Boston and Philadelphia were St. Louis 1913 (108-195, .356), Chicago 1948 (115-191, .376), Bay Area 1979 (12-199, .386), New York 1965 (127-197, .392), and Los Angeles 1992 (135-189, .417).

The best records are:



Four of these top ten featured a crosstown World Series, led by the 1906 victory by the White Sox over the Cubs; the others are St. Louis 1944, New York 1951 (Giants/Yankees as the Dodgers dropped the three-game NL playoff), and New York 1952 (this time Dodgers/Yankees). The banner years for the other cities were Boston 1915 (184-119, .607), Philadelphia 1913 (184-120, .605) and Los Angeles 2009 (192-132, .593).

The overall records for each city (for years in which they had multiple teams) are:



The cities in which the combined record has been good still have two teams; the ones in which they were poor do not. Shocking but true.

Thursday, April 21, 2011

Wayne Winston's Mathletics

The "book reviews" on this blog are almost always a day late and a dollar short. They are written and published long after the book, and my comments about them usually don't amount to a review but rather as a springboard from which to discuss other topics. This one is no different.

Wayne Winston is a professor of Decision Sciences at Indiana University's business school and a former consultant to the NBA's Dallas Mavericks. He published Mathletics in 2009 with the tagline "How Gamblers, Managers, and Sports Enthusiasts Use Mathematics in Baseball, Basketball, and Football."

If you are a regular reader of this blog or similar material, do not buy this book expecting to learn a lot of new things about sabermetrics. The sabermetric material is fairly standard, rudimentary type material--introductory-level discussion of run estimators, park factors, replacement level, the base/out table, win expectancy, and the like. I would also not recommend it to a novice, not because it is poor (there are elements I like and dislike, as I'll discuss below), but because there are better resources out there--internet primers, Bennett and Fluck's Curve Ball, and Lee Panas' Beyond Batting Average among others.

I am not particularly well-read on either football or basketball quantitative analysis, so I cannot definitively state the level of Winston's discussion on those topics. My guess is that the football discussion is fairly basic (with the caveat that football analysis as a field lags behind apbrmetrics), but that the basketball material is much stronger. It is certainly obvious from the writing that basketball is Winston's passion, and that the adjusted plus/minus ratings are a particular favorite.

Winston's writing is not particularly strong--he writes like someone whose favorite class was math (as do I). There are some minor slip-ups in the baseball discussion; these won't mislead the reader, but they also reflect the pedestrian nature of the material:

* Winston includes a formula for estimating batting outs that accounts for ROE by putting a multiplier on at bats. But this applies the adjustment to all at bats, including those in which we know a batter did not reach on an error (hits) and those in which the likelihood was very small (strikeouts).

* He refers to Keith Woolner's statistic as VORPP--Value Over Replacement Player Points. This makes sense in that he applies the replacement level concept to WPA points, but he also refers to Woolner's run based version as VORPP. Additionally, he credits the concept of replacement level to Woolner. In reality, Woolner did much to popularize replacement level, but the concept did not originate with him.

* Similarly, he credits the concept of park factors to Bill James. James had much to do with popularizing the notion that statistics could be corrected for park effect, but if any single person is to be credited with the concept, Pete Palmer would be an easy choice.

* There is a chapter that discusses player improvement over time by comparing annual performance, but it does so without even really addressing aging and survivor bias.

* The discussion of strategy is fairly bare-bones and deals only with basic estimates based on a standard run expectancy table.

There are positive things of similar magnitude to the list of negatives--for example, while he uses Runs Created, he explains that a theoretical team construct is necessary to make accurate player comparisons. As a whole, the baseball portion of the book is adequate without being excellent for a novice and a yawn for those well-versed in sabermetrics.

Being a novice myself when it comes to football and basketball analysis, I found the discussion in those chapters much more interesting. Focusing on a couple interesting football tidbits, Winston offers a version of the famed two-point conversion chart that incorporates the expected number of possessions remaining in the game. There is also a formula for the probability of a successful field goal in the NFL based on distance that I found interesting, although the model produces results that are clearly too high for very long kicks.

There is also a discussion of quarterback ratings, which have always interested me. Like every other sane person, Winston has little use for the NFL system, focusing his discussion on Berri's rating from Wages of Wins and his own adaptation of Brian Burke's regression of team categories against team wins. Isolating the categories from Burke's equation that can be related directly to individual quarterbacks, Winston offers the following as a quarterback rating:

1.543*(Yards - Sack Yards)/(Attempts + Sacks) - 50.0957*(Interceptions/Attempts)

If you factor out and ignore the 1.543 coefficient, and change the second quantity's denominator to (Attempts + Sacks), this can be rewritten as:

(Yards - Sack Yards - 32.47*Interceptions)/(Attempts + Sacks)

In this form, Winston's rating is very similar to a number of rating formulas, including the NEWS rating published by Bob Carroll, John Thorn, and Pete Palmer in The Hidden Game of Football:

NEWS = (Yards - Sack Yards - 45*Interceptions + 10*Touchdowns)/(Attempts + Sacks)

Breaking into editorial mode and stepping away from Mathletics for a moment, the treatment of a touchdown pass can be thought of as somewhat analogous to the sacrifice fly in baseball. The comparison is strained as touchdown pass is always a positive play from any perspective, while a sacrifice fly might actually reduce run expectancy.

A fairly large number of touchdown passes occur on short passes. Suppose a quarterback completes a three-yard touchdown pass. This will actually reduce his rating in Winston's ranking, as the quarterback's rating prior to the touchdown will be higher than three. By giving a positive weight to all passing touchdowns, one could ensure that a touchdown pass always increases ranking.

However, in doing so, one gives special treatment to the touchdown because it is a tracked category (like sacrifice flies). However, one could also track "sacrifice grounders" or "first down completions". These theoretical categories would also be cases in which a positive or somewhat positive outcome was achieved, but the statistics treat it as a negative (a batting out or a reduction of the passer's rating, assuming the completion was short). Giving special treatment to the recorded categories can thus be seen as unhelpful and biased by particular types of players that might be predisposed to one or the other.

Moving back to the book, most of my comments to this point have focused on the negatives. However, there are three things that Winston does really well:

1. Winston provides downloadable spreadsheets for many of the examples. This allows the reader to follow along with the work and to learn how to carry it out in Excel. Many of the Excel steps are explained in the text as well.

The drawback to this is that some of the why behind the math is glossed over in favor of a quick Excel solution. Winston's rating system for NBA and NFL teams basically boil down to finding the best-fitting solution for a system of linear equations to predict the point margin in each game. Winston doesn't explain the math in that manner, though, instead just explaining that the Excel solver is used to minimize error. While this gives the reader enough detail to produce their own ratings, and no one is actually going to solve hundreds of equations, I personally prefer a stronger emphasis on the underlying math.

2. The bibliography is excellent, as it includes not just a list of sources but descriptions of what they offer. For example, this is the description of Phil Birnbaum's Sabermetric Research blog:

This is perhaps the best mathletics blog on the Internet. Sabermetrician Phil Birnbaum gives his cogent review and analysis of the latest mathletics research in hockey, baseball, football, and basketball. This is a must-read that often gives you clear and accurate summaries of complex and long research papers.

3. Winston's description of Birnbaum's blog provides a nice transition into discussing the best thing about his approach. While Winston has excellent academic credentials (he is a professor of Decisions Sciences at Indiana and earned a PhD at Yale in Operations Research), but he does not beat you over the head with it. In fact, I don't think that his doctorate is ever explicitly referenced.

In any event, Winston mixes the research of other academics into his text, but he gives plenty of space to amateurs as well. Some academics that enter the sports arena seem to thumb their nose down at anyone who doesn't hold an advanced degree or a teaching position. Winston is not one of them. He even used one of Birnbaum's posts to offer a counterpoint to an academic paper on the NFL draft.

Winston's book provides a great example of how sabermetric knowledge generated by academics, amateurs, and everyone in between can be integrated, and how all parties can respect and learn from each other. It also gives analysts specializing in each sport a window into the work being done on other sports. Thanks to those attributes, Mathletics is a worthwhile read.

Wednesday, April 13, 2011

Comments on Baseball Prospectus 2011

At some point it becomes bad sport to write the same thing about an annual book--if there’s a certain characteristic of the book that you find yourself dissatisfied with several years running, it might be a you problem. It’s one thing to decide that a certain book is not for you; it’s another to continue to believe that it will when it’s obvious that the writers have something else in mind.

Much of what I could say about the Baseball Prospectus annual for 2011 is the same as I said about in 2010, and 2009…and so I’ll try to avoid saying it again. By now, it’s clear that BP is what it is, and that can either be a great thing or a bad thing or a mostly good thing, depending on your perspective. My perspective is that it’s mostly a good thing--the redeeming qualities of the book outweigh its flaws fairly easily from my perspective.

I still felt compelled to jot down a few comments on the book this year because I might have been a little unfair in nitpicking a few things in the past. Now that there is a lot of new blood on board, it’s more apparent that some of the issues (like stats not matching up between the comments and the data directly above) are systematic, and probably endemic to producing a book of this kind. To put together a tome of that size in a few months is a massive undertaking, and there are thousands of moving parts, so expecting them all to be dialed in to the same setting is unrealistic.

The cover still has the infamous phrase that I will not repeat about PECOTA; this is obviously out of the hands of the writers. They do redeem the cover with a great caption under the little photo of Albert Pujols.

That being said, I do have one major bone to pick with the new, slimmed down statistical offerings. It’s great that they stopped doubling up on metrics that measure the same thing (in the past, there have been simultaneous displays of VORP and WARP, or EqA and MLVr), and with one glaring exception the new stat lines still manage to give you most of the key metrics. That glaring exception is the lack of any kind of component ERA (or RA, which I’d prefer anyway) figure for pitchers.

It’s not simply a matter of limiting your choice to vanilla, while having to leave chocolate, strawberry, and cookies and cream aside (after all, there are a lot of flavors of component ERA). There is none whatsoever. Instead, BP has listed Fair RA, which is a fine metric constructed by Colin Wyers and the primary input for pitcher WARP. But if the choice is between having Fair RA and a component ERA in a book that is largely aimed toward predicting performance in 2011, it’s not a choice at all. Sticking with metrics under the BP umbrella, peripheral ERA and SIERA would fit the bill.

Of course, if I could strike any category from the pitcher stat line to clear space, it wouldn’t be Fair RA--it would be W-L or saves or WHIP. But since a big target audience for the book is fantasy players, that is not an option. However, it leaves everyone (including fantasy players) without a backwards looking metric that gives us the best estimation of how the pitcher’s overall effectiveness in the past. I certainly hope that they will figure out a way to include Peripheral ERA or SIERA or something similar in the 2012 edition.

PECOTA is in good hands with Colin Wyers, and I’m sure there are still some bugs to be worked out, so please take this comment as more amusement than criticism: some of the PECOTA comps seem way off. I’m sure this happened in the past, and I didn’t bother to make note of it, but two players that really stood out to me were Gregor Blanco and Nick Franklin. Blanco’s top comps are Richie Ashburn, Kenny Lofton and Freddy Guzman. One of these things is not like the other, and two of them are nothing like Gregor Blanco (Lofton was still in the process of breaking out, but had already established himself as clearly better). The Franklin comps are more understandable since he’s a younger player with less of a track record, but it’s still an odd juxtaposition to see a player ranked as the #44 prospect in MLB while his top comps are identified as Adrian Beltre (ok), Hank Aaron and Willie Mays.

There are only a few team entries that have extensive sabermetric (as opposed to applied sabermetric) content. One of these is the Arizona entry, and sadly I have a bone to pick with it. The author accepts the mainstream view that Arizona’s copious strikeout totals in recent campaigns had doomed their offense. He (or she; I still maintain it would be more interesting to know which author is responsible for the team entry) asserts that “when the majority of the lineup falls prey to empty at-bats of this sort, highly volatile run-scoring can result.”

While there have been some studies done on the relationship between shape of offense and scoring distribution, I am personally unaware of any comprehensive or well-established enough to make a statement like that without the need for supporting evidence. The only statistic brought in to support that position is that Arizona scored three or more runs per inning as much as the NL average, but scored two or less more often.

That is a very odd and not particularly helpful way to break down innings, because it lumps scoreless innings in with one and two run innings. To be absurd for a moment, if an offense never scored three or more runs an inning, and scored 0-2 in 100% of their innings, but 40% of those were one run and 10% were two runs, they would average a healthy 5.4 runs per game. It is true that Arizona scored in a smaller proportion of their innings than did the average NL offense--25.9% of Arizona innings resulted in a run scored compared to 26.5% for the league as a whole. But Arizona was more likely to have a multi-run inning (12.4%) than the average NL team (12.2%).

Another odd thing about this perspective is that it makes the inning the unit by which scoring volatility is measured. It’s true that the best perspective from which to understand how runs are scored is the inning level, since the events that transpire in each inning is independent of those that occurred in previous innings in terms of scoring in runs (I hope it’s clear that I’m talking about baserunners and outs from one inning affecting each other, not lineups turning over and pitchers being removed and the like, but you never can tell) but from a win/loss perspective, it is the run distribution per game that is crucial. Admittedly, the two are very closely related, but any time you extend the time period over which such volatility is projected, its impact is reduced.

One crude but simple and reasonably sensible way to consider the win value of a team’s per game scoring distribution is a method that I call Game Offensive Winning Percentage (gOW%) and have published here for the last three years. It is based on a Bill James idea; instead of estimating an OW% from average runs scored per game, use the team’s actual distribution of runs scored. If in a given season teams that score one run win 11.8% of the time (as they did in 2010), then credit the offense with .118 wins for each game in which they score exactly one run. Repeat for all scoring levels and average and you have an alternative OW%.

There are of course flaws with this method--the unit of games doesn’t always represent the same things (i.e. there are not always 27 outs per game), the use of the actual W% by runs scored in any given season is subject to sample size fluctuations, there is no adjustment for park, etc.--yet it’s still reasonable to think that if a team’s run distribution was particularly unusual, it would manifest itself in a comparison of gOW% to standard OW% based on average runs per game (in this case, without a park adjustment so as to better match gOW%).

The Diamondbacks led the NL in strikeouts in 2009 and 2010 and were second in 2008. In 2007, they ranked eleventh (and made the playoffs, see!), so those three seasons are the relevant high strikeout seasons for the team. In 2008, Arizona’s gOW% was .485 while their OW% was .479--considering their run distribution rather than just their average suggests an additional win. In 2009, it was .484/.483--no difference. In 2010, the split was .492/.502, which is -1.6 wins. So for the three years considered together, the net total is -.5 wins.

Of course, this does not conclusively demonstrate that Arizona’s offense was as efficient as a typical offense with their scoring average, and it certainly doesn’t allow us to make any statements about the effect of high strikeout offenses generally. However, neither does anything offered or referenced in the BP essay, yet the author chose to make much stronger assertions than I would dare to here.

My comments on strikeouts should not be taken as a negative judgment of the book as a whole--my book “reviews”, such as they are, generally serve as an opportunity to discuss issues raised by the author rather than to offer a summary judgment on the book itself. By now, you already know whether BP is a book for you or not.

Tuesday, March 01, 2011

Comments on Bill James Gold Mine 2010, pt. 2

2. Defensive Win Shares and Loss Shares

James has revamped Win Shares over the last couple of years to include Loss Shares. I think this is a very good thing, although I look forward to when (if?) the entire methodology is published. Without the full explanation, it's dangerous to comment about isolated details, but James' essay on "Explaining Defensive Win Shares to a Dead Sportswriter" is tough to ignore. My Twitter-friendly take on it: He's going to have trouble explaining it to a lot of people, not just dead sportswriters.

Again, it's impossible to evaluate the method while knowing so little about it, but James makes this extraordinary statement:

Making outs increases the team's responsibility to play defense. When you make more outs, that increases the team's responsibility to play defense. Therefore, if two players are the same in the field but of them makes more outs, the one who makes fewer outs has to come out ahead when you compare the player's defense contribution to his defensive responsibility.

Lest you think that was just a slip, he doubles down:

While we are in the habit of thinking of offense and defense in baseball as un-connected, they are in fact not un-connected. There is a very important connect between them, which is the rule that for every out you make on offense, you must record an out on defense.

Bill James is obviously a very intelligent man, and you a very intelligent reader, so I am hesitant to respond to this--the response should write itself. Limiting myself to a paragraph or less, I suppose it is technically true that each out on offense is matched by a defensive out, barring walkoffs and rainouts and the like. But there is no causation between the two. The rules of the game require three outs per inning and nine innings per team. Each team makes 27 outs regardless of the rate at which they use them (think OBA) or any other factor.

An individual who makes outs at a higher rate than some comparison player does not increase the number of outs that his defense must record. The defense must record 27 outs regardless of what an individual does at bat. What does happen is that by consuming excess outs, the individual batter leaves less outs to be consumed by the other eight members of his lineup, and fails to generate additional plate appearances for them.

James later seems to suggest that the revamped DWS-LS system assigns the same responsibility to field to each position, regardless of where it stands on the defensive spectrum. He then states his objection to offensive-based positional adjustments, and so it seems as if the stuff about making outs might be a backdoor way of applying positional adjustments. It's unclear, though, and still doesn't follow logically.

James’ discussion of positional adjustments also seems to gloss over the use of defense-based positional adjustments or the fact that most of us who still use offensive positional adjustments do so because we believe they provide a ballpark estimate of the defensive differences between the positions. When I use an offensive positional adjustment, I'm not saying that I think a shortstop with a 5 RG is a better hitter than a first baseman with a 5 RG. What I am saying is that the difference between aggregate offensive performance between shortstops and first baseman (when considered carefully and over a long period of time) approximates the inherent difference in defensive value.

You are certainly free to reject that argument (and many sabermetricians that I respect very much do just that), but please recognize that the sabermetrician using an OPADJ is likely not making the claim that a player's offensive contribution is altered by his fielding position.

More important than my own positional adjustment folly is an apparent failure by James to recognize that the positional adjustments that are now used most prominently in the community (generally Tango's, which have made their mark on the PADJs used in WAR figures from both Chone and Fangraphs) are based on estimates of the defensive difference between positions, sometimes informed by offensive averages. Furthermore, the sources do not lump the positional adjustment into the offensive ranking--they break everything (offense, fielding, baserunning, position, etc.) into smaller components, which are then summed to produce RAA, WAR, or some other total value metric.

Again, it is possible that I have misunderstood James' point, or that he has done a poor job of expressing himself, and that DWS is completely logical. However, I think it is going to take a much more thorough explanation of the system to give people that read the Gold Mine piece a lot of confidence in his methodology.

3. Strikeout rate

One of the most thought-provoking essays is "Whiff 7", which discusses the phenomenon of strikeout rates continuing to reach all-time highs. James argues that there is no end to this in sight under current conditions, as teams have an incentive to find power pitchers but no disincentive to find batters that avoid striking out. James argues that the standard deviation of power (he doesn't use that terminology) has decreased over time, and so league homer rates have gone up while the top individual performers hit about as many homers as they did in previous eras.

James then offers some suggestions of rule changes that would slow or reverse the trend. It's an interesting piece, and it didn't prod me to respond to it directly, but rather to make a tangential and mostly unrelated point about how we measure strikeout rates--a wholly unoriginal and stale one at that.

I have for a long time advocated using K/PA rather than K/IP as the measure of pitcher strikeout proficiency (I’m not claiming this is unique, as others have carried that banner with much more vigor and coherent arguments than I have offered). Through no effort of mine, the use of K/PA has increased in the sabermetric community, with sites like The Hardball Times and Fangraphs prominently utilizing K/PA.

As an example of how the different denominators can change perception, consider the point that most long-term successful pitchers have at least average strikeout rates. This is a point that the average fan still mystifyingly misses a great deal of the time. Take Greg Maddux for example. Maddux is apparently seen by some as a non-strikeout pitcher. Here is a table with his K/9 versus the league average, with KAA being strikeouts above average per inning:



For his career, Maddux struck out 6.1 per nine, while the league average was 6.4. He struck out 206 less batters than an average pitcher would have in the same number of innings. Without seeing the same figure for a lot of pitchers, it's hard to contextualize that, admittedly.

Suppose that instead you look at Maddux through K/PA:



Now Maddux' strikeout rate is essentially average--he struck out 17% of opposing batters, the same as the league average. Maddux' career rate is lower (it's actually 16.5% to 16.6%), but just barely so, and by this metric he only recorded 22 less strikeouts than average.

In Maddux' peak years (I think 1992-98 stand out), he was above-average even by K/9--+90 KAA, while he was an even more robust +196 when K/PA is the standard.

This is not intended to recast Maddux as a strikeout fiend--certainly he was not, even at his best. Still, Maddux' strikeout rate is more impressive when viewed in light of the number of opposing batters he actually faced rather than in terms of innings pitched, which really is just a measure of the percentage of outs a pitcher gets via the K rather (this is obscured by displaying strikeouts per 9 innings rather than strikeouts per 27 outs).

In addition to K/9, there are several other per-inning pitching ratios in common usage--H/9, W/9, HR/9, WHIP. What all of those have in common is that they are ratios of bad things (offensive successes) to good things (outs recorded). K/9 is a ratio of really good things (outs recorded by strikeout) to another set of plain old good things that includes the really good things (total outs recorded). As such, it's best viewed as a measure of a pitcher's reliance on strikeouts.

Monday, February 21, 2011

Comments on Bill James Gold Mine 2010, pt. 1

I quite enjoyed the third edition of the Bill James Gold Mine, even though I didn't get around to reading it until a few months after it was published. It jogged some thoughts, which lead to this post, which is not fully based on James' essays but on the semi-related paths they sent my mind down. To me, that is one of the tests of a really good sabermetric work--does it get you thinking, even if not about the exact topics covered? James' book passed that test for me.

However, I do think that the book would be stronger if it contained more of James' essays and less "statistical nuggets". The nuggets were of less interest to me, and seemed to be present in lesser quantity than they were in the first two editions of the book. The reverse was true for the essays, and those are what compel me to buy the book. Not being a subscriber to Bill James Online, I'm not positive about this, but I believe that James writes a number of additional essays in each year that are not included in the book.

If that is indeed the case, I believe that they'd be much better off to collect all of Bill's essays in the Gold Mine, and leave the nuggets for the individual to drudge up themselves online. Not only does the website lend itself more to the statistics (the data there is much more extensive than what can be printed in the book even if the book were the size of one of the old Great American Baseball Stat Books) and the essays to the printed page, but if there are any folks out there who still refuse to use the Internet and are interested in James, I'd think they'd be more enticed by the essays. A book of just the essays, with some other filler of some sort, would have a character not unlike that of the 1990-1992 Baseball Books, which I liked very much.

Of course, since it appears that the book is not even being published in 2011, those suggestions are for naught.

I have three subjects to touch on, two of which could be considered critiques and one of which is just a good old-fashioned tangent. This post went a lot longer than I originally intended so it's been broken up into two portions:

1. Starting pitcher rankings

The longest essay in the book deals with a system to rate starting pitchers based on where they place among other starters in each of their league seasons. James first ranks pitchers by Season Score (*), and then assigns points based on the pitcher's standing in the league. Each league season has 5.5 points per team available. In a fourteen team league, the top ranked pitcher gets 12 points, the #2 ranked pitcher gets 11 points, and so on down to the #11 pitcher who gets 2 points. There are also three-point bonuses, up to nine points per season, available for truly historic seasons. The resulting metric is called Strong Season points.

(*) James does not give the formula for Season Score in the article, but explains that is based on W, L, IP, ERA, K, W, and SV. "The point of the system is to evaluate a pitcher's record without context"..."This was a way of trying to say 'How good are the numbers themselves?', rather than 'How good was the pitcher who compiled these numbers?'".

Personally, I'm not sure that I have a whole lot of interest in rankings of pitchers based on a method that deliberately ignores context (and James certainly does not deny the importance of context). Setting my objection aside, though, it seems to me as if the Season Score is yet another result of a process that James has repeated over the course of his career: the re-invention of Approximate Value. Of all of his methods, my impression is that there is none that James personally likes more than AV. Even Win Shares is in some respects a return to AV--while it attempts to adjust for everything, it still expresses the result in an integer. The scale is higher than that of AV (a 20 AV would be an extraordinary season, while 20 WS is good but ordinary).

And so after attempting to adjust for everything, it seems James still had a void in his own toolkit, and so he filled it with the Season Score.

Digression aside, James found that a career total of 43 strong season points marks a fairly clear line for the Hall of Fame in retrospect. Only five pitchers retired for a significant length of time have more than 44 points and are not in the Hall--Vida Blue, Bert Blyleven, Ron Guidry, Carl Mays and Billy Pierce. James says that Blyleven and Guidry (60 points) are the only two pitchers that were far above 43 yet are excluded from Cooperstown. (Blyleven has been elected since James wrote the book and I wrote this post, obviously).

Since Guidry's is the most surprising result of James' survey, I'll take a closer look at him. I do not intend the discussion about Guidry to be a commentary on his Hall worthiness or even his value, but rather as a means of discussing the issue I have with the strong season method. It is important to note that James does not in any claim that the strong season method must be used in ranking pitchers, that it is better than any methods X, Y, and Z, or any such thing. James does not argue that Guidry should be in the Hall of Fame because of his showing in the system.

Guidry earned points for six seasons in James' analysis--1977-79, 1982-83, and 1985. Suppose we apply James' method, but use a different metric--a simple Runs Above Replacement, figured using total runs allowed and adjusted for park. How many points would Guidry earn under such a system?

* James ranked Guidry #6 in the AL in 1977, which is worth seven points. I have him #7, worth six points.

* James and I both have Guidry #1 in 1978 with an extraordinary season for 12 points (James awards the 9 point bonus, and I'll do so as well to keep things comparable). Guidry turned in 101 RAR, seventeen more than the next closest pitcher and nine more than any other AL pitcher in any of these six years.

* James had Guidry #3 in 1979 for ten points. I have him second, for eleven points.

* James ranks Guidry #11 in 1982 for two points. I have him all the way down at #26. His RA was 4.22 in a league in which 4.5 runs were scored per game, and he pitched in a moderate pitchers' park (.97 PF). At 34 RAR, he is eleven runs behind the eleventh-place pitcher (Geoff Zahn, 45). Presumably Season Score gives Guidry a boost because of his 14-8 record, one of the most impressive in the league (seventh in the league in Win Points).

* James ranks Guidry #4 in 1983 for nine points; I have him #6 for seven points.

* James ranks Guidry #2 in 1985 for eleven points; I have him #11 for two points. This is another season in which Guidry's W-L record seems to give him a huge season score boost (22-6).

Add it all up, and I have Guidry at 47 points--suddenly not that far above the Hall of Fame line James observed. I followed his scoring method exactly, but the results changed significantly simply by changing ranking methods.

More interesting, IMO, is how the use of in-season rank elevates the importance of very small performance differences. In 1979, Guidry ranked second in RAR at 71. However, Tommy John (71) and Jerry Koosman (70) were right behind him. Given that Guidry relied much less on his fielders, I strongly support the notion that he had a better season than the other lefties. Still, negligible differences in actual performance are given much greater impact when one uses a points system like James'.

Another example is 1985, in which Guidry ranks eleventh on my list at 61. Jimmy Key ranks sixth at 62--there are six pitchers within two RAR of each other. Guidry could very easily rank sixth in this season, which would be worth an additional five points. That would vault him from 47 points to 52 points, and give him a great deal more clearance over the HOF line.

This is not to say that James' ranking system is without its strong points with respect to its aims--it values peak performance and it sets an equal total value relative to the size of the league, which depending on one's perspective might be very good properties. My contention is that such a system is very sensitive to small changes in statistics, ones that would have no impact on a career-based evaluation. If Guidry had been evaluated at 62 RAR and thus sixth in 1985, the extra run saved would have zero impact on your evaluation of his career RAR total--and rightfully so. Allowing one run to exert a significant difference in a player's rank on an all-time list strikes me as utterly illogical and unsatisfactory.

You may object and say that I am using RAR rather than Season Score, and that Season Score is not subject to minute differences in performance having a large effect on rank order as is the case for RAR. While it is true that RAR and Season Score are very different methods, and that their application to Guidry might be very different as well, any metric is going to be subject to the same concerns when making a rank order over one season. There is always the potential that a very small margin could be the difference between a batting title and third place, between fifth in the league on a list and out of the top ten. That is true for any metric you want to pick, from BA to home runs to ERA to Season Score to RAR.

Monday, June 07, 2010

Comments on "Baseball Prospectus 2010"

For the second year in a row I have some thoughts to share on the BP annual, and I am not going to write them up in such a way as to feel comfortable calling it a "book review" (although I did give the post that label). It's not nearly formal enough, nor timely enough, nor extensive enough, to qualify for such a description. It is also a list of quibbles rather than a balanced look that praises the strengths of the book. I assume that anyone reading my blog is familiar with BP and doesn't need me to restate the table of contents.

Last year, I along with many others decried the lack of a player index in the book. This, happily, has been rectified, and now I'm just left to fume that Eric Fryer was not given a comment. On the other hand, the book contains no extra essays, for the first time since 1998 at least. Though the essays have been of uneven quality the last few editions, and have drifted away from sabermetric topics, they were always one of my favorite parts of the books. There have been some really good ones through the years, like Keith Woolner's piece on replacement level and Michael Wolverton's on pennants added.

Ditching the essays allows BP to devote more room to player comments (and an index!), and put essay content on their website, but I like baseball annuals that give you something to go back to in the years to come. In 2020, no one is really going to care what BP or anyone else thought of the Orioles or Yovani Gallardo in 2010 (let alone Yuniesky Betancourt). Essays about sabermetrics or the game in general are what give the Abstract or the Hardball Times Annual a shelf-life past June each year. However, I realize that BP is not attempting to be another Abstract (and their cover blurbs finally seem to recognize it too), so to some extent I am wishcasting the kind of book I prefer upon them.

The biggest quibble I have with the BP annual can be boiled down to confusion and contradiction with respect to the statistics used. Go to their website and take a gander at their glossary. It is next to impossible to use, it's disorganized, and it is largely devoid of formulas--even for metrics for which the formulas have been published previously.

The disarray of the glossary is mirrored by the haphazard use of statistics in the book. Written comments often diverge from the statistics listed just above. This is not new to BP 2010, which only makes it more annoying. Here are four ways in which the broad problem manifests itself, three of which I'll illustrate with examples drawn from the Reds team chapter alone:

1. Using both VORP and WARP, which both measure the same thing (value above replacement), with the major differences outside of the respective offensive metrics used being the unit (runs or wins) and fielding (ignored by VORP, included by WARP). This requires clarification of how players rank relative to replacement level, as in the case of Paul Janish:

As poor a hitter as Janish is, he's that good with the glove, enough so that he was able to keep his head above replacement level last year (note the difference between his VORP and WARP totals above, as WARP includes defense and VORP doesn't).

There is something to be said for displaying hitting and fielding value in separate columns when using an uber-metric due to the wider differences in estimates of fielding value between systems, but instead BP lists two separate metrics with two different scales.

2. Comments that use untranslated numbers, juxtaposed with the translated numbers just above them. See Willy Taveras:

Taveras's -14.3 VORP in 2009 was the third-worst in the majors.

Taveras's VORP is listed as -8.0 just seven lines above.

3. Using different metrics than the ones listed, that measure the same thing. See the Joey Votto comment:

Despite all that, August was his only poor month, and he finished the year fourth among all qualified major leaguers with a marginal lineup value per game of .397

MLV/G was previously listed by BP; it was taken out this year. The elimination of MLV/G cleaned up an issue with overall offensive rate duplication analogous to the doubling-up of VORP and WARP discussed above, as EqA is a measure of the same thing. There are significant methodological differences and significant unit differences, but unless there are special circumstances in which those differences are important, there's no need to use both (and simply pointing out that Votto was one of the most effective major league hitters is not one of those cases, as EqA surely concurred with that assessment.)

4. Ignoring BP fielding metrics in favor of other fielding metrics like UZR and Plus/Minus.

I think the authors are well within the mainstream of the sabermetric community if they trust UZR and Plus/Minus more than the BP fielding system, and they should be applauded for being willing to cite metrics published outside of BP. On the other hand, though, if you have so little faith in your own metric that you don't even want to quote it, why include it at all? Why pollute WARP with its presence?

The statistical introduction to the book is not immune from confusion. This year they do acknowledge that they now use Pythagenpat rather than Pythagenport, but the editor still missed a slip-up within just a couple of sentences. This was unintentional, but the issue is further confused by the use of the term generically without specifying whether what BP calls first, second, or third-order inputs are used.

I have to quibble with their choice of ERA-style metrics listed, although this complaint lies squarely in the realm of opinion and not methodological error. Included are Matt Swartz and Eric Seidman's new SIERA, which is based on batted ball inputs, and DERA, which adjusts actual ERA for team defense. That leaves an ERA estimator based on component statistics (H, W, HR, etc.) out of the annual for the first time in many years, as they previously included Peripheral ERA. I would like to see PERA included alongside SIERA, at the expense of DERA if necessary.

The most puzzling glossary comment comes on page xii:

We've transitioned this year to using a measure of VORP that is based on EqA; this has appeared for years on the BP Web site under the label of "RARP".

If true, I would heartily applaud this, as the old VORP is fueled by MVP, which is based on the flawed OBA*SLG*AB model of basic Runs Created. EqA is essentially linear weights-based, and thus a much more robust metric when applied to individuals. However, it doesn't seem as if BP contributors are at all clear on whether this change was actually made for the annual, and the VORP report on the website still appears to be MLV-based.

If in fact VORP listed in the annual is EqA-based, it makes the listing of both VORP and WARP even more curious, as they would be identical except for the inclusion of fielding and the conversion to wins. If both are going to be displayed, it would make all the sense in the world to display them in directly comparable units (be it runs or wins).

I am a fairly hard-core sabermetrician, and I am bewildered by the array of metrics used, so I can only assume that the average reader of BP has absolutely no idea what the difference between VORP and RARP is, let alone how to calculate them. While that may not be necessary to understand the implications of the results, it is necessary to enable people who are interested (like myself) to understand the differences.

Sometime after the book was published, BP changed the name of "Equivalent Average" to "True Average". The timing makes it seem as if this was a spur-of-the-moment choice, as the new annual presumably presents a perfect opportunity to change names. In any event, I'm not particularly fond of this change. I don't really care, as others do, that EqA has a long history--if something can be improved, I don't think there should be a statue of limitations, even when it comes to a name. I just don't think that "True Average" is a very good name.

For one thin, they abbreviate it as "TAv" (the other obvious choice would be "TA"). If I see a sabermetric stat with an abbreviation in that vein, the first one that pops into my head is Tom Boswell's Total Average. While EqA is obviously a better constructed and more useful metric than TA, TA has a much longer history, dating back about thirty years. While I have no problem with changing the name of one's own metric, I'm not crazy about doing so in a manner which potentially impedes on someone else's.

Secondly, "Equivalent Average" was a great name for what the statistic is. It is a measure of the rate of overall offensive production designed to look like an equally impressive batting average--an equivalent batting average. To call it "True Average" misses the mark for me on two counts:

1. A "True Average" should be expressed in meaningful baseball units--not units that have been assigned great meaning by custom (as in the case of BA), but units that actually have a great deal of meaning. Examples that work for me would be runs per out, runs per game, runs per PA, even OBA.

2. Batting average, for all its flaws, is straightforward. It truly is hits per at bat. To call another metric "True Average" implies (to me at least) that it is a truer measure of BA. Of course, it's not attempting to replace BA in the role of a measure of base hit frequency, but just in BA's role as a measure of offensive production.

Ultimately, the name of the statistic really doesn't matter, and this is admittedly a petty quibble. I still think "Equivalent Average" is a much better name.

Again, I'd like to reiterate that this is not a review of BP 2010 as a complete work, just a comment on the performance metrics employed, and the accompanying confusion. As usual, I'm glad I bought the book--I just wish I had a better understanding of where the numbers came from.