I am not a fan of the two wildcard playoff format, but my protestations were not considered and so it will soon be upon us. One thing I meant to do look at eventually was how the second wildcard would impact the probability of teams of a certain presumed strength winning in the playoffs. I’ve never gotten around to it, so I’m pretty much forced to look into it now or forever hold my peace.
The model that I will use to discuss this is admittedly simplified. It makes assumptions that are clearly simpler than reality:
* Teams have a constant strength from game to game. Even the strictest believer in baseball games as an expression of random chance disagrees with this as the identity of the starting pitcher obviously matters.
* Game outcomes are completely independent of one another. While game outcomes are largely independent, the argument for independence is weakened in the playoffs where decisions about how to manage the game (particularly pitcher usage) are clearly influenced by the status of the series.
* Home field advantage is uniform for all teams.
With these assumptions (along with others that have gone unstated), it is easy to construct a model of a playoff series. If one ignores home field advantage, the binomial distribution makes it very easy. That is the way I’ve approached these problems in the past, but I’ve decided to consider home field advantage and make the computations a tad more arduous this time (Of course, as you incorporate HFA into analysis of playoff series, you realize how inconsequential it is barring some intangible psychological force).
I have set up an (excessively) clumsy spreadsheet to do the math. The spreadsheet allows you to enter the playoff teams in order of seeding (i.e. A1 is the AL #1 seed down to N5, the NL’s second wildcard, with the National League assumed to have home field advantage--if the AL does, just enter the AL teams as N1-5 and the NL teams as A1-5) and enter a strength rating for each team (in the form of a win ratio as I use here). It then calculates the probability of each potential playoff series and their outcomes. The probability of each series outcome is figured on the other tabs, and if you are so inclined you could alter the home field pattern for each round. You can also enter the average W% for the home team. I’ve set this to .573, which is the World Series average for 1922-2008. The regular season average is usually around .540, so I think this is a fairly generous assumption in terms of strength of HFA.
The yellow cells are where the user should input custom data. I’ve not provided full documentation for every step as I doubt anyone will actually use this spreadsheet, but if you do and have any questions I will be happy to expound on the documentation. The spreadsheet can be accessed here (change html at the end to xls to download in Excel format).
In the post linked above, I looked at the winning percentages for all playoff teams and theoretical second wildcards for 1995-2010. I’m going to use these averages to set up a theoretical “typical” playoff scenario, and see how the probability of each team advancing to certain round varies with and without the second wildcard team. Using the actual W%s without adjustment to represent the strength of the teams is wrong for a couple of reasons, most notably that some regression is needed to estimate true quality and that no adjustment has been made for the unbalanced schedule, which is a big concern when performing interleague comparisons (and a smaller but still present concern for intraleague comparisons). However, exaggerating the differences in quality between teams will produce a liberal estimate of the differences between playoff formats, which may not be terrible for the sake of discussion. Also note that I’ve not made any adjustment for the fact that under the old format, the wildcard could be matched up with the #2 seed if the #1 seed came from their division.
Here are those average W%s and the resulting CTR (simply W%/(1 - W%) in this case) for each seed:

To rehash that earlier post, much of my antipathy towards the second wildcard and the fetishization of division titles is on display here. The AL wildcard has typically been one of the strongest playoff teams, while the wildcard in both leagues has a better average record than the third division winner. If MLB is hellbent on allowing a fifth team into the playoffs, then I would propose making the playoff between the two qualifying teams with the worst record rather than the two that failed to win their division. This complaint is water under the bridge at this point, though.
First, let’s run through the playoffs with a standard home field advantage (I’m using .543) and just one wildcard:

Now the same scenario, but with a special increased playoff HFA of .573. The last column is the marginal number of World Series victories per 1000 seasons relative to the .543 HFA assumption:

Remember, I’ve given the National League World Series HFA, so the NL teams get more of a boost from assuming a stronger HFA than do the AL teams. The NL picks up 9.6 World Series victories per 1000 seasons as a result of this stronger HFA assumption.
More relevant to the point of this post, here is the effect of adding the second wildcard to the mix. Again, let me emphasize that I’m assuming that there is no difference in expected W% from game-to-game. This is particularly relevant for the wildcard teams as one of the purported benefits of the extra playoff is that it will put the winner at a disadvantage entering the Division Series in terms of pitcher availability, since they will clearly have an incentive to use their best available pitching in the wildcard game. I’m not saying that you shouldn’t attempt to model this and other game-to-game factors when assessing playoff probabilities--but doing so complicates the exercise considerably. Instead, think of what I’m doing here as simply an analysis of the format itself rather than the consequences of that format upon the teams. These probabilities are based on the .573 HFA assumption. The last column is the marginal number of World Series victories per 1000 seasons relative to the single wildcard:

Keeping in mind that this analysis is not a true comparison of the previous format to the current one as it doesn’t account for the wildcard being unable to face a divisional opponent in the Division Series, this actually makes me feel a little better about the two wildcard formats, as it increases the probability of the best teams (#1 and #2 seeds, although the AL wildcard is generally in that class as well) winning the World Series. #1 seeds get an easier Division Series matchup by getting the second wildcard roughly 40% of the time. Obviously the #5 seed benefits the most, going from out of the picture to being a long shot. Equally obviously, the first wildcards take a huge hit.
The interesting result is the decrease in W% for the #3 seed, the reason for which is not immediately obvious. The cause is the increased likelihood of facing the #1 seed in the LCS (and the increased likelihood of facing the other league’s best teams in the World Series).
Still, the effects on the division winners’ odds are relatively small, not much different than the difference in assuming that the typical player HFA is .03 wins greater. The brunt of the impact is felt by the first wildcard.
Tuesday, September 25, 2012
Playoff Probabilities
Tuesday, September 11, 2012
Meanderings
* There is always some grumbling about September roster expansion, and the supposed ills it inflicts on the game, but it seems to have reached a fever pitch in September 2012. There are a lot of calls for some kind of reform, whether it involves severely curtailing the practice or (most popularly) forcing a manager to declare 25 active players at the start of every game.
I don’t have a problem with roster expansion, myself--my preference would be to keep the status quo. I will also admit to not having read all of the pieces that have been written about this, so it is quite possible that someone has prominently beaten me to the punch with the following suggestion. Rather than having the manager declare a 25 man active roster, why not simply limit the manager to using 25 players in a particular game?
Such a rule would give a manager in-game flexibility that would be absent in the case of a pre-declared scratch list. In a close game, he would be free to use extra players in situational roles. In a blowout, he would be able to use his mop up relievers and get young players into the game. But he would not be able to use any more players than he could during the rest of the season.
There are a couple of drawbacks to this rule that come to mind. One is that 25 players is still an increase over the number that is typically deployed during the rest of the season, even in fairly unusual cases. The four excess starting pitchers usually are excluded from game action, especially in the American League. I’d argue that this is a good thing--it allows for some additional substitutions while preventing abuse, but if one is concerned about any change in the behavior of managers, it’s a valid criticism.
The other is a logistical issue rather than a baseball issue, but it would be a little harder to keep track of. Announcers would be completely bewildered as a team approached the substitution limit, and while the omnipresent lineup cards should be sufficient for managers and umpires to keep up, it’s not hard to imagine some confusion arising.
* Everyone has an opinion on Stephen Strasburg. I don’t, really--I certainly agree with the principle of being cautious with pitchers, particularly young pitchers, those who have prior injury histories, and those of extraordinary talent--all three of which fit Strasburg. What I do have an opinion about is the intestinal fortitude of Mike Rizzo and anyone else who took responsibility for the final decision. I consider it a pretty bold stance to take, given that there is almost no outcome in which they do not receive heavy criticism.
If Washington fails to win the World Series, the question of how they might have performed with Strasburg will be raised incessantly. Even a quick sweep in the Division Series in which one win would have not stemmed the tide will not get them off the hook, because psychological factors will be raised (“Strasburg could have won game one and completely changed the momentum”, “The Nationals players would have been more confident with Strasburg available”, etc.) And if they lose a seven game World Series in which Edwin Jackson gets roughed up a couple of times--well, if I was Rizzo, I’d consider hiring a food taster at that point. So far I’ve only discussed the 2012 on-field consequences. A bigger outcry will come if Strasburg gets hurt again, particularly if it’s in the next two years.
What’s remarkable about this decision is the near certainty with which it will be judged as a failure by mainstream observers. Perhaps my imagination is too limited, or my faith in sound reasoning on behalf of mainstream observers artificially low, but it’s difficult for me to imagine a scenario in which the shutdown is considered to be a success. The odds are very good that Washington will not win it all, with or without Strasburg. Steven Strasburg is quite unlikely to have an injury-free career. The takeaway for me is that Rizzo must really believe he’s made the right call.
* I am an unabashed supporter of the World Baseball Classic, so I’m taking it as my duty to update you about the upcoming qualifying tournaments. The existence of these has largely gone unremarked upon.
For the first time, four spots in the sixteen team field will be up for grabs. The twelve countries that won games in the 2009 WBC are automatic qualifiers (Japan, Korea, China, United States, Mexico, Italy, Netherlands, Dominican Republic, Venezuela, Cuba, Australia, and Puerto Rico). The first two qualifiers open up next week.
The qualifiers are four-team double elimination tournaments, the format of which will be familiar to those who watched the 2009 WBC or follow the NCAA Tournament. In Jupiter, FL, South Africa will play Israel and Spain will play France. In Regensburg, Germany, Canada will play Great Britain and Germany will play the Czech Republic. In November, the other two spots will be decided. In Panama City, Panama will play Brazil and Colombia will play Nicaragua. In Taipei City, the Philippines will play Thailand and Taiwan will play New Zealand.
Handicapping these tournaments is silly (after all, the Netherlands beat the Dominican Republic twice in the 2009 WBC). Speaking broadly without knowledge of the actual makeup of the teams, there are clear favorites in the Germany and Taiwan qualifiers, as Canada and Taiwan are much stronger baseball nations than the second-tier European and Asian countries, respectively.
The other two are fairly wide open. WBC rules allow people who could be but are not citizens of a country to play, which allows Israel access to Jewish players. Israel will be managed by Brad Ausmus, and faces a very weak field, so they may be the favorite. Panama played in the first two WBCs without much success. Colombia and Nicaragua have both produced their share of major leaguers, and even Brazil now has their own big leaguer in Yan Gomes.
Just to provide a general sense of how the participating countries have performed in recent international tournaments, here is each country’s rank in the IBAF World Rankings divided by qualifier. (These rankings don’t do justice to countries with strong baseball that don’t field teams in many international tournaments such as the Dominican Republic, and obviously are based on tournament results and don’t tell one anything about the quality of WBC team being fielded. Nonetheless, it’s kind of fun to look at--and it’s even more fun to look at the full ranking list, which includes countries such as Bolivia, Myanmar, and New Caledonia, which I had to look up on Wikipedia. Also note that the top-ranked non-participant, Netherlands Antilles, is included with the Netherlands for the WBC):

It is also worth noting that the IBAF website states that it will now bestow the title of world champion on the WBC winner rather than the soon to be defunct World Cup winner.
Monday, August 27, 2012
Sabermetric Generations
Allow me the indulgence of going meta-sabermetric here. I don't really like doing that, as this is the point at which you are writing about sabermetrics itself rather than baseball, and baseball is a heckuva lot more interesting than sabermetrics. However, I think there are some things about the field that those of us who fancy ourselves as sabermetricians should consider. This post is a half-developed missive on some of those things.
My basic premise here is that sabermetric people can be divided into three generations--not perfectly, of course, but as a general classification. I say "people" because I don't want to make it about who is or is not a "sabermetrician" (Do you have to publish your own run estimator to be a sabermetrician? Write a blog? Do any research at all?), but just about people who would consider themselves to be either practitioners or consumers of sabermetric research, or both.
These three generations are not defined strictly by age, but rather by when you came of age as a saberperson (or how, but the when and the how elements are very closely related). My three groups are:
1. Pioneers--This is by far the most restrictive group, as it only includes those who actually did pioneering sabermetric research (whether they called it by that name or not). Earnshaw Cook, George Lindsey, Pete Palmer, Bill James, and the like are the pioneers.
2. Second Wave--These are the folks who came to sabermetrics largely through the work of the pioneers--reading the Baseball Abstract or The Hidden Game or various SABR publications or The Diamond Appraised, and the like. They may or may not have gone on to become researchers themselves; they may just be consumers of research. It is also possible that their own inquisitiveness led them to sabermetrics without a firm push from Bill James or another pioneer, but they still came onto the scene after the work of the pioneers had been published. Many, many people fall into this group, and even listing a few would be foolish. I consider myself in this group although I also share a few traits with the third group, as I explained before.
3. Internet generation--These are people who have come to sabermetrics in the last 10-15 years and may have done so without ever reading the work of the pioneers. Their interest in sabermetrics is young enough to have been fueled by reading the work of second wavers (or of course their own inquisitiveness). A typical path to sabermetrics for a member of the Internet generation would have been to read Rob Neyer. From there, they sought out BP or Bill James.
Before I use these classifications to make a point, I need to issue a couple of disclaimers. The first is that I am not an evangelist for sabermetrics. I don't go to the airport and hand out flowers, or go to people's doors and hand them tracts. I don't really care if you are interested in sabermetrics or not, and I don't tailor my writing to appeal to folks who are on the fence.
So when I express my concern about something below, it's not borne out of any fear of what overzealous internet posters will do the reputation of sabermetrics or anything like that; it is simply out of an ordinary desire for intelligent and factual discourse.
The second disclaimer is that this is not a get off my lawn post. As I stated in the piece linked above, I straddle the fence between the second wave and the internet generation, and while I might wish to place myself in the former group, you could make a reasonable case that I am in fact a member of the latter group. No group is inherently better or worse than any other; this post is about certain negative traits of some members of the internet generation, but there are many positive things that can be said about the internet generation.
The third disclaimer is that this whole matter of putting saberpeople into groups and then describing those groups is obviously dangerous for the same reason that forming any sort of artificial groups of people is. So it should go without saying that I am not claiming that all members of a generation share certain characteristics or behave exactly the same way.
My concern is about the fact that with the wealth of information available today, particularly through sites like Baseball-Reference, Fangraphs, and StatCorner, it has become quite possible for members of the internet generation of saberpeople to cite statistics without really understanding them at all. Of course, second wavers also had this ability, but it wasn't so instantaneous. You had to wait for new books to be published in the spring or you had to go through the trouble of figuring statistics yourself.
Of course, there are many members of the internet generation who do fantastic research, develop their own statistics, or figure stats themselves. They are not who I am talking about. I am talking about the subset of folks who don't do their own research, don't endeavor to truly understand what the numbers mean, and yet still talk about them authoritatively.
With sabermetrics prominent on the internet baseball scene, it is much easier to learn about the field. This is on the whole a very welcome improvement, but it is also easier for people to get indoctrinated into sabermetric principles without fully understanding them. There is a sabermetric-brand of conventional wisdom which can be just as misleading as the conventional wisdom of the traditionalists when it is wielded by an individual who has not done his own legwork, but has simply read it or told it and believed it to be true.
The ubiquity of sabermetric ideas in online discussions of baseball makes it quite possible for members of the internet generation to be introduced to sabermetrics almost simultaneously to the moment at which they become serious baseball fans. This pretty much happened to me as recounted in the earlier linked post, although it was primarily through printed works of pioneers rather than through the internet. That being the case, I can't criticize this path of discovery. However, I do think that there might be a critical thinking advantage to having first accepted the conventional wisdom, gradually looking at it skeptically, and then having those questions reinforced by the discovery of sabermetrics. Today it is quite feasible for young baseball fans to skip the conventional wisdom altogether and jump right into sabermetric ideas.
A specific example of the kind of thing I'm talking about is the notion that any pitching statistic like ERA that does not incorporate DIPS principles is worthless, and that only FIP or tRA or a similar metric is appropriate. The problem is not with the truth that ERA has a lot of biases which have often been overlooked, or that FIP is a better predictor of future performance. The problem comes when a relatively good measure like ERA (and one that measures actual runs allowed, which are unquestionably important if not wholly attributable to the pitcher) is thrown in the dustbin as if it is no more useful or telling than Batting Average or raw RBI count, and that anyone who even considers it is classified as a dinosaur.
In defending ERA, I am not saying that it is inherently wrong to come to the conclusion that FIP or Metric XYZ is not the best tool for measuring pitcher performance--just that it is wrong to reflexively come to that conclusion, and that sometimes the advocates of such a position can take on a zealous tone. This tone often echoes that of reflexive anti-sabermetric screeds.
You may be thinking to yourself that I am attacking a strawman, and that no one actually thinks like that. I didn't quote/link anyone specifically, because singling out message board posters is not the point of this discussion--but sentiments of this sort are out there.
This is where my admission that this post is half-developed becomes painfully clear, because I really don't have anything to offer about what can or should be done about this (given that I am not a sabermetric evangelist, my developed answer would most likely be centered around the premise that individuals are responsible for their own rhetoric). At this point, it will devolve into a paean to do-it-yourself sabermetrics.
There are now at least four large-scale implementations of WAR floating out there--Rally/Baseball-Reference, Fangraphs, BP's WARP, and the Baseball Gauge's WAR. As a result, people will sometimes wonder why sabermetrics can't have a meeting of the minds, hash out all of the differences, and produce one unified version of WAR that can be presented to the wider world.
There are a number of reasons why I think this is a bad idea (the danger of presenting any one metric as *the* uberstat; the legitimate uses of alternate baselines, run estimators, park factors, league adjustments, position adjustments, and all of the other components that go into WAR; the notion that consensus is a positive for its own sake), but that's irrelevant, because it will never happen. The hypothetical moment that it did happen would prove Gary Huckaby right--sabermetrics would be dead. Enforcing standardization for problems in which the answer is often subjective would discourage innovation and discredit alternate views on the question of ability v. value (among others).
Additionally, the existence of one version of WAR, widely accepted and presumably published by major websites, would in my estimation do more to discourage people from figuring their own statistics than anything else in the history of the field--more that Total Baseball or Baseball-Reference or Fangraphs. If you were a new convert to sabermetric thinking, and were told that there was one metric that was the best and was freely available, what would be your most likely reaction?
1) Awesome, this sabermetrics thing is not nearly as complicated as all of the screeds suggested it would be.
2) Darnit, I was hoping to find the unified theory of everything myself.
3) Darnit, I can't believe I don't get to try to figure out how to use a slide rule, or make a spreadsheet, or do SQL coding, or however these saberwhatevers get their results.
There is a lot of value in figuring your own sabermetric statistics, even by just applying other people's metrics. There's no better way to learn about the inputs and how they are combined than by actually walking through the process yourself.
At one time, in order to have up-to-date access to sabermetric stats, you pretty much had no choice but to figure them yourself. This resulted in a waste of a lot of man-hours of sabermetricians, but it also meant a group of people that were better informed about the construction and therefore the objectives and philosophy of the metrics they were using. The easy availability of statistics today is a great boon to the field, to be sure, and it is particularly great for serious practitioners who already have a good understanding of metric construction. I am not a luddite--the downside of a potentially less-informed average consumer of sabermetrics does not outweigh the benefits--but it also should serve as an impetus for transparency in computational explanations and for continuing reminders of the why in addition to the numerical results themselves.
Monday, August 20, 2012
Ballpark Thoughts
My apologies to anyone who reads the title of this post and expects to read a discussion of some arcane aspect of estimating park factors. On Saturday I visited Great American Ball Park in Cincinnati for the first time for the Cubs/Reds doubleheader. I am by no means a well-traveled fan when it comes to attending MLB stadiums--GABP raises my lifetime count to five (Jacobs Field, Municipal Stadium, PNC Park, Tropicana Field). What follows are some random thoughts (and snark at the expense of my southern neighbors):
* I expected to be underwhelmed by the ballpark. I can’t point to any particular influence, but I thought that the general consensus on GABP was that it was not on par with the best of the new parks.
Having only been to four of the current parks, I can’t place GABP on a grand continuum, but given my lowered expectations, I was quite impressed. For the day game, I purposefully bought a ticket in the very last row of the stadium behind home plate. My main motivation was shade, but it offered a great opportunity to take in the park from the bird’s eye view. For the second game, I sat in the right field moon deck, which was useful because it provided the opposite orientation.
The Ohio River and the hills of Kentucky provide nice scenery for those facing the outfield. While the park is not situated close enough to the field to allow for splash hits, the river is certainly quite prominent in the view past right field, and the lack of a second or third deck in right field leaves the scenery largely unimpeded.
Those looking towards left field have a much less interesting view--the left field bleachers and the basketball arena block any view. The impediment of the arena validates the decision to leave right field open, as otherwise you wouldn't be able to see much beyond the park from any perspective.
From the outfield perspective, you mostly just have a view of the stadium. Most of the skyline is obscured, with the top of the (surprise) Great American building the highlight. The PNC “power stacks” are a bit of an annoyance from these seats--not because they block the view (they’re behind you) but because you can feel the heat from the napalm or whatever exactly it is that they shoot off. This would be a plus at a cold April game, but is annoying otherwise.
Of course, all this talk about the view misses the primary point of visiting a stadium, which is to watch the ballgame and not the scenery.
* The silliness of the attempt to manufacture a “game day experience” is by no means unique to GABP nor to MLB, but in my limited experience Cincinnati takes the cake. One of the most annoying features added at Jacobs Field over the last few seasons are way too cheery “hosts” who appear on the scoreboard, going around the park and telling you about all the exciting and fun things to do at the park. GABP has these as well, so I was edified about the pre-game concert featuring a band doing a particularly bland cover of the Black Crowes’ version of “Hard to Handle”, about the Big Red Machine exhibit at the Reds Hall of Fame (the Reds “dominated baseball in the 70s”...I’m sure the A’s really felt dominated), and various other distractions.
The Reds also feature an extraordinary number of mascots. There are four. One is Gapper, the generic fuzzy monster that almost every team has a version of. The next is Mr. Red, a giant baseball head who manages to look much more menacing than Mr. Met. Then there is Mr. Redlegs, the old time giant baseball head (you can tell by his taste in mustaches) who appears to cheat at the mascot race (only conducted on the scoreboard, which is a plus). Finally, there is Rosie the Red, easily the most creepy mascot in history.
Prior to the day game (the night game was not as bloated), there were three first pitches (a logical conundrum), a kid yelling “play ball”, an honorary captain, and a delivery of the official game ball to the mound. Many traditional religious services have less ceremony.
Again, none of this is unique to Cincinnati, but I’ve previously been fortunate to not encounter so much of it at once.
* For the night game, the first 20,000 fans received a 1995 replica hat with Barry Larkin’s #11 on the side. The good people of Cincinnati really wanted to ensure that they received their hats. The lines to get into the ballpark before the gate opened were very long. I wasn’t quite ready to go into the park yet (I like getting there early, but ninety minutes is a little much for me), but I bowed to the inevitable and got in line to ensure that I too would receive a 1995 Reds hat.
* Two things struck me about the area surrounding the ballpark. The first was the pandhandlers. Now maybe I don’t go to the right (wrong) places, but in the two cities I am most familiar with (Cleveland and Columbus), the panhandlers are not nearly as sophisticated. All of the Cincinnati pandhandlers stand there with cardboard signs displaying their sob stories.
The other is scalpers, or more precisely, the lack thereof. Apparently, Cincinnati has fairly strict (and thus asinine) laws against selling tickets above face value. Surely this is still happening, but it is clearly done more discretely than it is around Jacobs Field or Nationwide Arena.
* Along the river, there are a set of columns which have plaques on each side. These comprise the Steamboat Hall of Fame. Spoiler alert: Many of the steamboats in the Steamboat Hall of Fame met unpleasant demises.
* Charging admission to your team’s Hall of Fame is almost as pretentious as pretending that your team was established in 1869 when it was actually established in 1882.
Monday, August 13, 2012
Standing Still
2012 marked the second season of Greg Beals’ tenure as OSU baseball coach. It was not an encouraging season for the future of the program, and I’ve never been more perplexed by on-field strategy, which is saying something for college baseball. This is not to say that I’m declaring Beals to be incapable of restoring the program to the excellence it enjoyed throughout most of Bob Todd’s tenure, but it looks like it will be a long road from here.
The Buckeyes overall record did improve slightly, from 25-26 to 33-27. But the difference was wrapped up in non-conference play, as OSU’s Big Ten record was essentially unchanged (12-11 to 13-13). OSU played a much less ambitious early season schedule, playing five games against southern teams (Georgia Tech and Coastal Carolina) but the rest against northern opponents. In fairness, OSU’s ISR (Boyd Nation’s ranking system) did rise to #94 from #156.
In conference play, OSU was consistently mediocre. OSU took one of three in series against Purdue, MSU, Nebraska, Illinois, and PSU. The other three series were sweeps--two at home in Ohio’s favor against Minnesota and Northwestern, and one to Indiana on the road. The wipeout in Bloomington came in the season’s final weekend and dropped the Bucks into a three-way tie for the sixth and final seed in the Big Ten Tournament. While the Buckeyes won the tiebreaker, it certainly felt as if they had backed into it.
In the Tournament (the last in a four-year arrangement to hold the event at Huntington Park), the Buckeyes rallied to beat Penn State, then lost to #1 seed Purdue. Once in the losers bracket, they beat Nebraska to stay alive but had their season ended by MSU.
OSU’s .550 W% ranked fourth in the Big Ten (Purdue led at .763) and fifth in EW% with a similar .548 (Purdue led at .732). In PW%, OSU looked a little worse (.520, fifth) with Purdue sweeping the W% flavors at .723. The Buckeyes ranked in the middle of the pack in both runs scored (5th at 5.52) and runs allowed (6th at 5.00).
OSU’s offense did one thing well, something that is close to my heart--draw walks. Their .140 W/AB ratio led the conference, well above the average of .100 and far above Illinois’ .107 which ranked second. In fact, Big Ten walk rates were tightly clustered, with the other ten teams ranging from just .089-.107. OSU’s ranked in the middle of the pack in batting average (.269 versus a .278 average) and isolated power (.086 versus the .100 average). The Bucks tacked on the conference’s most productive stolen base effort, leading the conference with 86 steals against 27 caught. This fact was surprising to me for reasons I’ll expand upon below.
Catcher remained a rough spot for OSU, as junior Greg Solomon posted a 38/6 K/W ratio and .252/.283/.396 line. Freshman Aaron Gretz showed a terrific eye (19 walks in 91 at bats), but little else (.253/.382/.286). First baseman Josh Dezse repeated as one of the team’s most productive hitters (second on the team at 11 RAA), but his power remained an enigma. Dezse tied the school record by belting three homers in a game at Georgia Tech, but hit just two for the remainder of the season. His .120 ISO represented a twenty point drop from his freshman season.
Second baseman Ryan Cypret had a nightmarish campaign a year after being one of the team’s most productive hitters, slumping to .236/.337/.304. Third baseman Brad Hallberg turned in a terrific senior campaign, leading the team with 12 RAA on the strength of a .311/.400/.431 line. Sophomore transfer Kirby Pellant represented an upgrade over OSU’s 2011 shortstop production, but at .274/.358/.340 was below average (-2 RAA).
Freshman Pat Porter took over the left field job as the season progressed, and compiling a pretty average .266/.360/.322 (if you are noticing a pattern, this team had a lot of middling averages and high walk rates with minimal power). Sophomore Tim Wetzel was actually the team’s third-most productive hitter by RAA (+6) thanks to his team leading OBA (.403), but no thanks to his lack of power (.056 ISO for a .336 SLG). David Corna, the primary right fielder, had a rough senior season (.241/.317/.390). Sophomore transfer Mike Carroll (.279/.360/.333 in 186 PA) and Joe Ciamacco (.291/.342/.330 in 111 PA) filled out most of the remaining playing time in the outfield corners and at DH.
Only two other players got significant playing time. Senior Brad Hutton served as part-time DH against left-handed pitchers, managing an average RG thanks to his walks (.220/.350/.340 in 60 PA). Freshman Ryan Leffel served as the utility infielder and could be OSU’s third baseman in 2013. For 2012, though, he could have been called “Josh Dezse’s glove”, as his main role was taking over third base when Dezse moved from first base to the mound (with Hallberg moving from third to first). Leffel appeared in 39 games, but only 3 were starts, and he was limited to 28 PA.
OSU’s fielding (admittedly these metrics leave a lot to be desired) was unremarkable, matching the conference average with a .941 mFA with a .674 DER versus the average of .677.
Before the season, OSU’s weekend starters were expected to be junior Brett McKinney, lefty JUCO transfer Brian King, and sophomore Greg Greve. But McKinney and Greve pitched poorly and lost their spots, with sophomore transfer Jaron Long emerging as the staff ace. Long made three relief appearances before establishing himself as the #1, and was the only starter to turn in an above average performance (+17 RAA). Long is a finesse righty who works in the high eighties at best, relying on his control (just 1.2 W/9).
King slotted in as the #2 starter, a bit of a disappointment given the hype with which he arrived. King was solidly average with -1 RAA. The #3 spot remained in flux until midway through the Big Ten season, when sophomore transfer John Kuchno earned the job. Kuchno was not particularly effective (-7 RAA), but his size and arm made him an eighteenth round pick of the Pirates, with whom he signed.
Greve (-3 RAA in 50 innings) and McKinney (-3 RAA in 71 innings) served as midweek starters and will again vie for the rotation in 2013. The bullpen was anchored by Dezse, who was very effective (2.86 RA, +7 RAA in just 28 innings for seven saves). His strikeout rate (6.0 K/9) continues to lag behind his stuff. Beals only had one lefty with experience in the pen, so senior Andrew Armstrong led the team with 36 appearances spanning just 28 innings. Unfortunately, Armstrong was not nearly as effective as in ’11, his 6.75 RA driven by 26 walks in just 28 innings. Junior sidearmer David Fathalikhani was effective, +5 RAA over 29 innings as the Armstrongs’ matchup counterpart. Freshman Trace Dempsey is being groomed as Fath’s replacement, but was not effective in his freshman campaign (5.63 RA in 32 innings).
In a second season of observing Greg Beals as coach, I have become absolutely mystified by the man’s strategy. Beals has increased OSU’s reliance on the bunt and basestealing. OSU’s ratio of sacrifices to (singles + walks) was .06 in 2012 and .07 in 2011, compared to .03, .05, .03 in Todd’s last three seasons. Beals called for many more steals this season as well, which worked out well--OSU led the Big Ten in steals with a solid percentage (75).
However, it was Beals’ fascination with one particular stolen base play that really gets my blood boiling. Beals is obsessed with the delayed steal of home with 2 outs, runners at the corners. Beals surely dreams about this play every night. I wish I had an easy way of counting how many times this was attempted, but my rough guess is once per series. It rarely worked; it might be insulting to call it a high school-level play. It was especially absurd to keep trotting it out in Big Ten play, as if the other coaches in the conference were a bunch of rubes with no ability to scout and no institutional memory.
OSU will be a popular pick to compete for the Big Ten title in 2013. The only key players who were lost to graduation/draft are third baseman Hallberg, right fielder Corna, starter Kuchno, and reliever Armstrong. The incoming freshman class was not ravaged by draft signings as Beals’ 2012 group was, and figures to infuse some pitching options. But my observation (anecdotal only) is that some of the most overrated teams in college sports are mediocre teams that return a lot of starters. The Buckeyes lack power, they lack quality starting pitching outside of Long, and to date they lack a coach who has proven that he can assemble a Big Ten contender.
Monday, August 06, 2012
All Models Are Wrong
The statistician George Box once wrote that “Essentially, all models are wrong, but some are useful.” Whether the context in which this was written is identical or even particularly close to the sabermetric issues I’m going to touch on isn’t really the point. Perfect models only occur when severe constraints can be imposed.
Since you can’t have a perfect model, the designer and user must decide what level of accuracy is acceptable for the purposes for which the model will be used. This is a question on which the designers and evaluators of sabermetric methods and the end users often disagree. In my stuff (I’ll use “stuff” because “research” is too pretentious and “work” is too sterile), my checklist of ideal properties would go something like this:
1. Does the method work under normal conditions? (i.e. does the run estimator make accurate predictions for teams at normal major league levels of offense)--This is the first hurdle that any sabermetric method must clear. If you can’t even do this, then your model truly is useless.
2. Does the method work for unusual conditions?--I'm going to draw a distinction here between “unusual” and “extreme”. Unusual conditions are those that represent the tails of the usual distributions we observe in baseball. A 70-92 team is not unusual, but a 54-108 team is (and keep in mind that individuals will exhibit a wider range of performance than teams). If the method fails under unusual conditions, it may still be useful, but extreme caution has to be taken when using it. Teams and players that threaten to break it will occur often enough that one will quickly tire of repeating the caveats.
3. Does the method work for extreme conditions?--Extreme conditions are the most unusual of cases, ones that will never occur at high levels of baseball over large samples. A player that hits a home run in every at bat will obviously never exist (although a player could have a game in which he hits a home run in every at bat). A method that can’t handle these extremes can still be quite useful. However, a method that can handle the extremes is much more likely to be a faithful representation of the underlying process. Furthermore, methods that are not accurate at extremes must begin to break down somewhere along the way, so if a model truly works better at the extremes, there’s a good chance it will also produce marginally better results than an alternative model for some of the unusual cases which will actually be observed.
And simply as a matter of advancing our knowledge of baseball, a model that works at the extremes helps us understand the true relationship between the inputs and outputs. Take W% estimators as an example. We know that, as a rule of thumb, 10 runs = 1 win. This rule of thumb holds very well for teams in the usual range of run differentials in run-of-the-mill major league scoring conditions. But why does it work? Anyone with a spreadsheet can run a regression and demonstrate the rule of thumb, but a model that works at the extremes can demonstrate why it is true and how the relationship differs as we move away from those typical conditions.
4. How simple is the method?--I differ from many other people in the relative weight given to the question of simplicity. There are some people who would rank it #1 or #2 on their own list, and would be willing to accept a much less accurate estimator if it was also much simpler.
Obviously, if two approaches are equally accurate (or very close to being equally accurate), it makes sense to use the simpler one. But I’m not a fan of sacrificing any accuracy if I don’t have to. Limits to this principle are much more relevant in fields in which more advanced models are used. However, most sabermetric models (particularly for the type of questions that I’ve always focused on) really are not complex at all and do not tax computer systems or trigger any other practical constraints on complexity. People might say that Base Runs is more complex than Runs Created, but the differences between common sabermetric methods are on the level of adding one more step or one more operation.
Now that I’ve attempted to define where I am coming from on this general question, the specific trigger for this post was Aroldis Chapman’s FIP. Trent Rosencrans of CBS pointed out that Chapman’s FIP for July was negative. It goes without saying that this is a breakdown in the model, as a negative number of runs scored makes no sense.
I have always been prone to overreact to a few comments on a site, and run off and compose a blog to respond not to a strawman, but to an extreme minority opinion. That’s probably the case here.
Still, it is interesting which metrics and which results set off this sort of reaction, and which generally don’t. A number of us tried unsuccessfully for years to argue for the replacement of Runs Created with Base Runs, but a number of very intelligent people resisted the idea. Arguments were advanced that Base Runs was too complex or that given the errors inherent to the exercise, Runs Created was good enough.
OPS+ also remains in use as a go to quick and dirty stat, due largely to its prominence on Baseball-Reference. Yet OPS+ clearly undervalues OBA and can also, in extreme situations, return a negative number of implied runs. ERA+ distorts the true difference in runs allowed rate and makes errors in aggregation across pitchers and seasons seem natural, but B-R had to backtrack quickly when they replaced it with something better because of the backlash.
It should be noted that some of the reaction to any flaws in FIP are likely related to general disagreement and distrust of DIPS theory as a whole (Colin Wyers pointed this possibility out to me). DIPS has always engendered a somewhat visceral reaction and remains the most controversial piece of the standard sabermetric toolkit. An obviously flawed result from the most ubiquitous member of the DIPS family is the perfect opportunity to lash out.
It’s no mystery why FIP fails for Chapman’s July. FIP is based on linear weights, and any linear weight estimator of absolute runs will return a negative estimate at low levels of offense. For example, a game in which there are three singles, one double, one walk, and 27 outs will result in a run estimator of around -.1 runs. [3*.5 + 1*.8 + 1*.3 - 27*.1] Run scoring is not a linear process, but there are many advantages to using a linear approximation (I won’t recap those here but this should be familiar territory). However, if the weights are tailored for a normal environment, and become less reliable as the environment becomes more extreme.
In July, Chapman’s line looked like this:
IP H HR W K 14.1 6 0 2 31
There is no linear run estimator that will be accurate for a normal context that will survive something this extreme. Fortunately, it is an extreme context observed over a very small sample size. A month may superficially seem like a significant split, but fourteen innings is roughly equivalent to two starts.
Going back to my checklist for a metric above, I do place a high weight on accuracy for extremes. While I am personally comfortable with the use of FIP for quick and dirty situation, I happen to agree with some of the critics that it isn’t particularly appropriate for use in a WAR calculation (presuming that you want to include a defense-independent metric as the input for WAR at all). It doesn’t make a lot of sense to sweat the small stuff as WAR does, except for the main driver (FIP). While the issue of negative runs will not be present over the long haul, using a linear run estimator for individual pitchers is needlessly imprecise by my criteria.
There are at least a couple different Base Runs-based DIPS formulas out there--Voros McCracken himself has one), I have one (see “dRA” here), and there could be others that have slipped my mind. Using dRA, Chapman’s July checks in at .68, which seems pretty reasonable.
The moral of the story is that our methods will always be flawed in one manner or another. Sometimes the designer or user has a tradeoff to make between simplicity and theoretical accuracy. Depending on the question the metric is being used to answer, there may be a reason to change one’s priorities in choosing a metric. Ideally, one should be as consistent as possible in making those choices. At the risk of painting with a broad brush, it is that consistency that appears to be lacking in some of the reaction in this case.
Tuesday, July 24, 2012
On Run Distributions, pt. 7: W% Estimates
There are a number of things one can do with a per game run distribution algorithm. One can look at the scoring patterns of teams to see if they are more or less efficient than a typical team with their average runs scored per game. While this can be done in reference to the empirical distribution for a given set of teams, having a distribution formula that works at multiple points allows one to do some customization, like accounting for park factor, that is problematic when working with the empirical data.
Another application that’s near and dear to my heart is using the run distribution tool to fuel a winning percentage estimator. Most win estimators you see are expressed in terms of a relatively simple function of runs scored and runs allowed, and you never really see how the sausage is made. But the way games are won is to score more runs than the other team. Essentially, any W% estimator is just trying to approximate the proportion of games in which a team will score more runs than it allows. With a function for run distribution, we can write this out explicitly.
Let Ps(k) be the probability of scoring k runs in a game, and Pa(m) the probability of allowing m runs. Then the W% of a team can be estimated as:
W% = Ps(1)*Pa(0) + Ps(2)*[Pa(0) + Pa(1)] + Ps(3)*[Pa(0) + Pa(1) + Pa(2)] + Ps(4)*[Pa(0) + Pa(1) + Pa(2) + Pa(3)] + ...
Of course, there are a number of other ways we could express the same idea, but while the equation above may not be the simplest, I believe it is the clearest. If you score 0 runs, you cannot win. If you score 1 run, you win if you allow 0 runs. If you score 2 runs, you win if you allow either 0 or 1 runs.
There is a complication that arises when employing this logic with a discrete distribution however (if we had a continuous distribution, we would be integrating the difference between the runs scored and runs allowed curves rather than summing, but the idea is the same.) With a continuous distribution, the probability of achieving any given k runs scored or m runs allowed is 0, so one need not be concerned with the run distribution method setting k equal to m. With the discrete distribution, though, we need some way of handling a situation in which a team is estimated to score and allow the same number of runs.
For the sake of this issue, I’ll assume that the Enby distribution is modeling runs in the first nine innings of the game, and that if the game is tied, it will go to extra innings. What is the probability of extra innings?
P(Extra Innings) = Ps(0)*Pa(0) + Ps(1)*Pa(1) + Ps(2)*Pa(2) + Ps(3)*Pa(3) + ...
The probability of extra innings is the probability that you score the same number of runs as you allow. Simple enough. Incidentally, if we assume that the distribution of runs scored and allowed are identical (like we might if we considered the league as an entire unit with some uniform R/G rather than as 30 individual units that only average that R/G when combined), the equation becomes:
P(Extra Innings) = P(0)^2 + P(1)^2 + P(2)^2 + P(3)^2 + P(4)^2 + ...
Knowing the percentage of extra inning games is not enough to estimate W%--we also need to know the percentage of those contests that the team goes on to win. For this, we will fall back once again on the Tango Distribution, which allows us to consider scoring on the inning level.
Before I actually use the Tango Distribution, allow me to walk through this with a simple example. The important factor in determining the probability of winning in extra innings is the percentage of the time that a team scores more runs in an inning than they allow. As soon as you do that, you win. As soon as you allow more than you score in an inning, you lose. As long as you score as many runs as you allow in an inning, the game continues.
Let’s suppose that a team has a 25% chance of scoring outscoring their opponent in a given inning and a 20% chance of being outscored. This means that there is a 55% chance of a tie in each inning, which extends the game. Thus, the probability of winning eventually is:
.25 + .55*.25 + .55^2*.25 + .55^3*.25 + .55^4*.25 + ...
You have a 25% chance of winning in the tenth inning and a 55% chance of their being an eleventh inning. In each subsequent inning, there is also a 25% chance of winning and a 55% chance of the game continuing. This expression can be solved as follows:
.25[1 + .55 + .55^2 + .55^3 + .55^4 + ...] = .25*1/(1 - .55) = .5556
However, you don’t even need to deal with the 55% probability of an additional inning, thanks to the Craps Principle. As you can see, the expression .25*1/(1 - .55) simplifies to .25/.45 which is equal to the probability of winning the first round divided by the sum of the probability of winning the first round and the probability of losing the first round, so we can just take .25/(.25 + .2) = .5556.
Bringing the Tango Distribution back into play, let Fs(k) be the probability of scoring k runs in an inning and Fa(m) the probability of allowing m runs in an inning. By “win inning”, I mean outscoring the opponent in a single inning; by “lose inning”, I mean being outscored in a single inning.
P(win inning) = Fs(1)*Fa(0) + Fs(2)*[Fa(0) + Fa(1)] + Fs(3)*[Fa(0) + Fa(1) + Fa(2)] + Fs(4)*[Fa(0) + Fa(1) + Fa(2) + Fa(3)] + ...
P(lose inning) = Fa(1)*Fs(0) + Fa(2)*[Fs(0) + Fs(1)] + Fa(3)*[Fs(0) + Fs(1) + Fs(2)] + Fa(4)*[Fs(0) + Fs(1) + Fs(2) + Fs(3)] + ...
P(win in extra innings) = P(win inning)/[P(win inning) + P(lose inning)]
P(extra innings) = Ps(0)*Pa(0) + Ps(1)*Pa(1) + Ps(2)*Pa(2) + Ps(3)*Pa(3) + ...
P(win in 9 innings) = Ps(1)*Pa(0) + Ps(2)*[Pa(0) + Pa(1)] + Ps(3)*[Pa(0) + Pa(1) + Pa(2)] + Ps(4)*[Pa(0) + Pa(1) + Pa(2) + Pa(3)] + ...
W% = P(win in 9 innings) + P(extra innings)*P(win in extra innings)
Let me demonstrate how this estimate is figured with a table. Let’s take a team that averages 5 runs scored and 4 runs allowed per game. We first use Enby to estimate the probability of scoring or allowing k runs (I’ve capped scoring at 25 runs here; technically, you need to go to infinity):

The columns “score” and “allow” are the probabilities of scoring k runs. “allow <” is the probability of allowing less than k runs. “win 9” is the probability of winning the game in nine innings while scoring k runs, and is equal to the probability of scoring k runs times the probability of allowing less than k runs. “extra” is the probability of extra innings with a score of k-k, and is the product of the probability of scoring k runs and the probability of allowing k runs. The sum of “win 9” is the probability of winning in 9 innings, which works out to .5409 in this case. The sum of “extra” is the probability of extra innings, which is 9.95%.
To complete our analysis, we need to do a similar analysis on the inning level using the Tango Distribution:

Here, I’ve added a “score <” column, which is the probability of scoring less than k runs. “lose” is the probability of losing given k runs scored and is the probability of allowing k runs times the probability of scoring less than k runs. “tie” here is the same as “extra” in the above chart, although it is also 1 - win - lose. The sum of “win” is the probability of winning the inning; the sum of “loss” is the probability of losing the winning. The probability of winning the inning divided by the sum of the probability of winning and losing the inning is the probability of winning an extra inning game (per the craps principles). In this case, the probability of winning an inning is 24.8%, the probability of losing an inning is 20%, and the probability of tying an inning is 55.2%.
Thus, the probability of winning given extra innings is .248/(.248 + .20) = .5534, and the overall probability of winning is:
.5409 + .0995*.5534 = .5960
Remember, .5409 is the probability of winning in 9 innings; the probability of winning given that the game only goes nine innings is .5409/(1 - .0995) = .6007. Obviously the team with the advantage has a greater probability of winning a nine inning game than winning a one inning game.
For comparison, Pythagenpat with z = .28 estimates that a 5 R/4 RA team will have a .6018 W%, a difference of .94 games over a 162 game season from the Enby estimate. The Tango-Ben distribution (using the same methodology as what I just demonstrated except substituting the Tango-Ben estimate of the game scoring probabilities for Enby) estimates .5953 when using c = .767, which is a good match for the Enby estimate. However, Tango found that for applications involving two teams in a head-to-head matchup, a c parameter of .852 produces better results. Using .852, the Tango-Ben distribution estimates a .6011 W%, a much better match for Pythagenpat.
It is quite possible that Enby would benefit from a separate set of parameters for use in a head-to-head matchup. This could be accomplished by modifying the variance targeted by the Enby parameters, and when I pick this topic back up in a few months I have some ideas on how to do that.
Given the complexity of the W% estimate and its questionable accuracy (at least with the current default parameterization) I have not endeavored to carry out more elaborate tests. For now, it will sit as an intellectual exercise rather than as a method I use.
Monday, July 16, 2012
On Run Distributions, pt. 6: Series Review
This post won’t introduce anything new--instead I’m just going to summarize what I’ve already done, giving you a full example of how to calculate the Enby distribution estimates for a given R/G level. I’ll also provide a spreadsheet with the parameters for each .05 R/G increment between 3-7 so that you don’t have to do all these calculations yourself.
Let’s suppose we have a team that averages exactly 5 R/G (in fact, there is such a team in my sample data--the 1984 Red Sox), and we’d like to estimate their game-level scoring distribution using the Enby distribution methodology. The first step is to estimate the variance of their runs scored per game:
Step 1: Estimate the variance of runs scored per game.
Variance = 1.43*(R/G) + .1345*(R/G)^2 = 1.43*5 + .1345*5^2 = 10.5125
Step 2: Use the mean and variance to estimate the parameters (r and B) of the negative binomial distribution (these formulas are equivalent to what I’ve presented before as explained below):
B = .1345*(R/G) + .43 = .1345*5 + .43 = 1.1025
r = (R/G)/(.1345*(R/G) + .43) = 5/(.1345*5 + .43) = 4.5351
Step 3: Use the negative binomial distribution to estimate the probability of scoring 0 runs:
q(0) = (1 + B)^(-r) = (1 + 1.1025)^(-4.5351) = .0344 (call this value a for ease later on)
Step 4: Use the Tango Distribution to estimate the probability of being shutout, which is equal to the Enby distribution (zero-modified negative binomial) parameter z:
RI = (R/G)/9 = 5/9 = .5556
z = (RI/(RI + .767*RI^2))^9 = .0410
Step 5: Using your spreadsheet, use trial and error (or a solver if you have that that level of functionality) to estimate a new value of r. In choosing this value, you need to ensure that the average R/G predicted by the Enby distribution equals your sample R/G (5 in this case). This needs to be done simultaneously; use the following formula to estimate the initial probability:
q(k) = (r)(r + 1)(r + 2)(r + 3)...(r + k - 1)*B^k/(k!*(1 + B)^(r + k)) for k >=1
Then modify it as follows:
p(0) = z
p(k) = (1 - z)*q(k)/(1 - a)for k >=1
The mean is calculated:
p(1) + 2*p(2) + 3*p(3) + 4*p(4) + ...
The new value of r is the value that, when used in conjunction with this methodology and the previously calculated values for B and z, produce a mean equal to the desired R/G (5 in this case, with a corresponding r of 4.571.
So we have determined that the Enby distribution for a team that scores 5 R/G has parameters (B = 1.1025, r = 4.571, z = .041). The formulas for p(0) and p(k) calculate the probability of scoring k runs in a game.
How does our plot for the 1984 Red Sox compare to their actual scoring output?

Of course, we don’t expect a great fit for every team-season. Even if we assumed that there were no variations in run distribution due to the characteristics of an offense, the 162 game sample size would cause deviation from the expected values.
I have calculated the three parameters at each interval of .05 R/G between 3 and 7. While we have some reason to believe that the Enby may be semi-accurate outside of normal ranges, I’m not going to recommend its usage outside of the scoring range of normal teams. Getting a lot more precise than .05 is probably overkill as well, but given my limitation in having to solve for r by trial and error, I’m also limiting the gradients as a matter of practicality.
Here is a link to the spreadsheet. Enter your R/G (only values between 3 and 7 are supported) in the shaded yellow cell. The spreadsheet will round this to the nearest .05 for you. P(k) is the probability of scoring k runs in the game, r is for computation purposes (it is the product of r*(r + 1)*(r + 2)... as applicable), and nb is the probability from the normal distribution without zero modification.
Since I now have a table with the parameters over the 3-7 R/G range, it would feel inappropriate not to make scatterplots and look for patterns. First, z against R/G:

The red line is the z values; the thin black line is an exponential regression line that is a decent match for the data over this range. z is the parameter that needs the least investigation, though, as it is calculated via a formula based on the Tango Distribution. The formula makes sense, and there’s no mystery about why it works. The regression equation is superfluous and will certainly fail at low levels of R/G (it will predict that a team that averages 0 R/G will only be shutout in 12.66% of games).
Here is B against R/G:

B is a linear function of R/G. This is also not a surprise. Remember that B = variance/mean - 1, but I’m estimating the variance as a function of the mean. In fact, B can be simplified to B = .1345*(R/G) + .43, keeping in mind that I have used a fairly crude estimator of the variance, which is an area that might well be improved upon.
The parameter for which behavior is not defined by a formula is r:

Over this range, r is almost linear as a function of R/G. It can be modeled very closely over this range by a quadratic regression. I wouldn’t want to assume that a function can be used to estimate r consistently over a wider range of R/G, and even if it did, I wouldn’t want to advocate it as the value of r should be chosen to ensure that the expected R/G equals the actual R/G. In any event, it’s interesting to see how the parameters might behave in relation to R/G.
This post is running a little shorter than most of the others, so I’ll throw in something that would have gone in the odds and ends post that will close this series. For the last few years I’ve been looking at runs scored and allowed distributions at the end of each season, and in that time the most interesting team I’ve seen is the 2011 Red Sox. Boston led the majors in runs scored, but based on the empirical W% by runs scored in the majors in 2011, their actual distribution of runs scored would have led to an estimated 6.2 less wins than one would assume from just looking at their average runs scored. I thought it would be interesting to look at such a team again with the Enby distribution.
Boston averaged 5.4 R/G, which from the table above means their run distribution will be estimated as Enby(B = 1.1563, r = 4.7, z = .0331). Graphing their actual distribution against the expected, we get this:

The Red Sox were shutout a lot more than we’d expect, and while we expected the mode of their runs per game to be 4, they actually scored 4 runs in 18.5% of their games compared to an expectation of 12.8%. They also were clearly below expectation in games of 5-8 runs scored, which are games that a team has a very good chance of winning. The distribution skewed more to right than expected, games in which gaudy runs scored totals have much less of an impact on wins as the marginal value of each run is quite low.
Another way to visualize this (and as you can tell from this series, charts aren’t really my thing--I'm using a lot here, but not to great effect and only because I think tables of numbers would bore you and require more exposition) is to graph the cumulative percentage of team runs scored as we progressively add in games in which k runs were scored.
Boston was shutout eleven times; obviously those games contributed zero runs. They scored one run twelve times, which contributed 12 runs. They scored 875 runs overall, so this represented 1.37% of their total output. They scored two runs fifteen times, for a total of 30 runs. So games with 0-2 runs represented 42/875 = 4.8% of their total output. Continuing in this vein, we can get a graphical sense of the share of their runs that came on the tails of the distribution:

I’ve included the Enby distribution expectation as well as the overall 2011 major league average on the graph. The average major league team tallied 88.6% of its runs in games in which ten or less runs were scored, while we’d expect a team that averages 5.4 R/G to have scored 80% of its runs in such games. However, Boston only scored 71.7% of its runs in those contests.