Monday, July 27, 2009

An Unusual League

Disclaimer: This post doesn't really have any direction; it wanders around to no real end. It also seemed a lot more interesting when it was in my head than it did after it was on the screen.

As you know, between 1900 and 1960, each major league without exception was comprised of eight teams, with the team with the best regular season record taking the pennant. The expression "first division", still in limited but diluted use today, referred to the top four teams in the circuit. Over the course of those six decades, there was one league that clearly stood out from the others in terms of the gap between the first division and the second division.

If one endeavors to quantify the amount of balance in a league's standings, there are a number of different possible approaches. The statistically-minded among us might immediately think about measuring the standard deviation of team winning percentage in each league-season, for instance.

But what about measures that would specifically measure the discrepancy in quality between the first division and the second division? Of course there are a number of different routes that one could go, but among the most obvious simple approaches are:

1. the fifth place team's games behind the pennant winner--this will tell you how large the gap was between the top of the first division and the top of the second division. I'll call this GB(5)

2. the winning percentage of the fourth place team--this is the target you're shooting for if you want to be a first division team. I'll call it W%(4).

3. the winning percentage of the fifth place team--this tells you where the second division begins in terms of wins. I'll call it W%(5).

4. the fifth place team's games behind the fourth place team--this is the direct gap between the first and second divisions. I'll call this GB(4-5)

5. the aggregate winning percentage of the first division--or of the second division, but it doesn't matter, as these two mathematically must be complements. I'll call this FD%.

These three indicators are all related to the quality gap between the first and second divisions, but come at it from slightly different perspectives. The first four approaches both ignore the entirety of the second division except for the fifth place team, but by doing so they establish the boundary between the two divisions. The last combines each group, but does nothing to temper the influence of outliers (on the high or low ends).

You could of course come up with other measures, but this is not intended to be a rigorous statistical examination--I just want to be able to establish that the league-season in question was somewhat unusual, and these rudimentary measures are sufficient for that purpose.

This is not a trivia post--if you want to guess, do it now, because the league in question is the 1950 AL. Here are the standings:


The imbalance, perfectly divided into two groups, jumps right off the page. The Yankees captured their second straight pennant by three games over the Tigers, with the Red Sox and Indians also in the hunt. But the second division lagged far behind, with the Browns and A's losing the equivalent of 100 games in a 162 game schedule.

You certainly don't need to formally look at any data to know that these are odd standings. Oddities like this are what can make flipping through a baseball encyclopedia so rewarding for those who dabble in statistics. There is a silly but tangible sense of discovery when you find some unusual statistical line or set of standings that you had not been previously aware of.

Anyway, this is a sabermetric blog, so you're going to be stuck with some pseudo-analysis rather than just a "gee whiz!". First let's look at the standard deviation of W%, which as mentioned above speaks to the balance of wins in the league but not specifically to the chasm between the first and second divisions. Still, out of the 123 major league-seasons in the 1900-1960 period, the 1950 AL ranks 14th in standard deviation of W%:


Many of the highest standard deviations occurred in the first twenty years of the period, so the 1950 AL ranks second in the post-war, pre-expansion era. Still, it doesn't stand out as anything remarkable in this respect as just four years later the standard deviation would be greater in the junior circuit. Here are the 1954 standings:


In this case, the chasm was between third and fourth place rather than fourth and fifth, and the standard deviation was high largely due to the Indians' rampage coupled with a strong effort from the Yankees (their 103 wins would have won the AL pennant in any other season between 1947-1960).

Moving on to the measures that specifically address the difference between the first and second division, we first have GB(5), which tells us how close the top of the second division was from the pennant. The 1950 AL was well above average but not remarkable in this regard, ranking in a tie for sixteenth-highest with the 1934 AL:


Of course GB(5) is strongly related to the performance of the pennant winner. You can see from the table that many of the league-seasons featured the great teams of the period...the 1906-07 Cubs, 1954 Indians, 1927 Yankees, 1931 A's, and the like. So I also looked at GB(5) divided by wins of the pennant winner, which bumps the 1950 AL up to thirteenth place.

Next we have W%(4), the floor of the first division, and this is where the unique nature of the 1950 AL starts to shine through. Cleveland, in fourth place at 92-62 (.597), had the highest W% of any fourth place finisher of the period, and it wasn't even close:


As you can see, it was by no means the first time that the Indians had a standout record for a fourth-place finisher.

In terms of W%(5), the ceiling of the second division, the 1950 AL made the bottom three (lowest W% by a fifth-place team):


Here it's the 1931 AL that leads the way, with St. Louis topping the second division with a 63-91 mark. As you might imagine, the race for fifth was very close, with Boston just one game back and Detroit two.

So it is not surprise that when we look at the gap between the first and second divisions, no league managed to come within five games of the 1950 AL:


Finally, we have FD%, which is the aggregate W% of the first division clubs. The 1950 AL comes in sixth:


As you can see, the 1950 AL leads the way among all post-1932 leagues, with the aforementioned 1954 AL next in line.

I hope that the combination of the "look test" and the data above will be enough to demonstrate that the 1950 AL was a uniquely two-tiered circuit. For those of you who write good history articles, I think that a brief history of how the AL franchises fortunes ebbed and flowed so as to create the conditions necessary for this historic imbalance would be a very interesting piece.

I don't write good history articles, so the Cliff's Notes (and potentially misleading summary) of the second division could be as follows:

* The White Sox never really recovered from the Black Sox scandal, with eight games back in 1940 the closest they got to a pennant.

* The Senators were solid contenders in the mid-20s and early 30s, winning three pennants, but outside of that, there's a reason "First in war, first in peace, last in the American League" was in use.

* The Browns had to share St. Louis with the Cardinals, and while neither team was strong in the first twenty years of the century, the Cards blew by them in on-field success by the late-20s and became the more popular draw, despite Bill Veeck's desperate efforts to win the patronage of the city's fans.

* Connie Mack was never able to rebuild the A's all the way again after selling off his stars of the 1930s--they had been respectable in 1947-49, but 1950 saw them collapse.

What would make such a piece more interesting is the fact that the form of 1950 generally held throughout the rest of the pre-expansion period. The degree of polarization between the strong and weak teams was not nearly as strong, of course, but with one major exception, the teams essentially stayed in their divisions throughout the decade.

The table below gives the finish for each franchise (sticking with their 1950 abbreviation in the case of Philadelphia (Kansas City) and St. Louis (Baltimore)); the first division finishes have been bolded:


As you can see, the Yankees stayed in the first division for the next ten years, while the Indians missed just once and the Red Sox thrice. The Senators stayed in the second division the whole time, with the A's escaping just once and the Browns franchise just once.

There was one significant change from the standings of 1950, and that was the reversal of fortune for the Tigers and White Sox. The Tigers would make it back into the first division just twice over the period, while the White Sox joined the Yankees in never missing over the next ten years.

Not only did the 1950 AL feature a huge gap between the first and second divisions, the second division also represented a sort of permanent league underclass and the first division a permanent group of contenders, with the aforementioned exception of the White Sox and Tigers respectively (of course, the large gap does indicate that the second division teams had a lot of work to do, so this is not entirely surprising). The second division saw three of its four members move to greener pastures, while the teams of the first division all remain in the same place today. With the exception of the White Sox, it would be fifteen years before a second division team of 1950 was able to win a pennant (the '65 Senators (Twins)).

After that, things got better quickly for the underclass, as the Orioles would emerge as the most consistent AL franchise of the next twenty years and the A's, after another move, would become just the second franchise to win three consecutive World Series. Meanwhile, the first division Indians tumbled into thirty years of hopelessness, emulating the historical examples of the Browns and Senators. But the particular state of imbalance in the AL, best demonstrated in the standings of 1950, had held for a long time.

Tuesday, July 21, 2009

On the World Series Home Field Advantage

A week ago, the American League once again defeated the Neanderthal League (*) in the All-Star Game, securing home field advantage for the World Series. The "This time it counts" mantra about the game is premised on the notion that home field advantage is a significant thing to have (or at least the hope that TV viewers will believe that it is). So it is only natural to look back through history and see how home teams have fared in the World Series.

Let's start off with some theoretical calculations based on a few assumptions. Assume that the two teams are evenly matched, that there is no home field advantage, and that the outcome of each game is independent of any other. Therefore, each team has a 50% chance to win each game, and we can calculate the expected frequency of a 4, 5, 6, or 7 game series using the geometric distribution (I apologize for this digression as many of you know this better than I do):

P(x+r game series for one team) = C(x + r - 1, x)*(1 - p)^x*p^r

Where x = number of failures before r successes, p = probability of success, and C is the combination function

In this case, our successes are victories by the eventual series winner (always r = 4), x is losses by the eventual series loser (0-3), and p = .5.

C(x + r -1, x) is the number of different of distinct sets of wins and losses that can occur in the series. C(3, 0) is used for a four-game series, and is equal to 1--the only string of wins and losses that can produce a four-game series is WWWW. The formula for combinations is:

C(n, x) = n!/(x!(n-x)!)

So C(4, 1), the number of different combinations that can produce a five-game series, is 4!/(1!(4-1)!) = 4*3*2*1/(1*(3*2*1)) = 4. You can confirm this, as there are four possible strings (LWWWW, WLWWW, WWLWW, and WWWLW) that produce a five-game series. In fact, you can logically work out all the combinations fairly easily without the math for this application since we are only dealing with a seven-game series.

In a five-game series, the fifth game must be a win (same for the sixth and seventh games of six and seven-game series, respectively). So the victor can lose game 1, game 2, game 3, or game 4.

In a six-game series, the victor can lose games:
12, 13, 14, 15, 23, 24, 25, 34, 35, 45 = 10 combinations

And in a seven-game series:
123, 124, 125, 126, 134, 135, 136, 145, 146, 156, 234, 235, 236, 245, 246, 256, 345, 346, 356, 456 = 20 combinations

Anyway, doing all the math (and then doubling since we have only considered this from the perspective of one team), the theoretical probability of a given series length is:
4 = 12.5%
5 = 25%
6 = 31.25%
7 = 31.25%

So theoretically (since WS home field sites are on a 12-345-56 pattern), in 43.75% of World Series, the number of home games will be equal. 25% of the time, the team with the home field disadvantage on paper will actually play more home games, and 31.25% of the time the team with home field advantage on paper will get to benefit from it--if and only if there is a game seven.

We'll get back to some theoretical stuff later, but let's look at the actual empirical World Series results. I considered all World Series from 1922-2008 (1922 is when the seven-game series returned permanently) with the following exceptions:

* 1922 and 1923--both Giants/Yankees series, in 1922 they shared the Polo Grounds, and in 1923 they didn't follow the 12-345-67 pattern
* 1943-45--in the war years, a 123-4567 format was used to cut down on travel (and in 1944, the Cardinals and Browns shared Sportsman's Park, which would have made it unusual in any case)

First, let's look at the empirical proportions of series by length:


As you can see, the empirical and theoretical don't actually track particularly well. I'm not going to discuss this phenomenon in-depth here, but it is something to keep in mind when we delve back into theoretical stuff at the end of the post. The assumptions are all faulty to some degree or another--the teams are not evenly matched, the results of the games are not truly independent (Even if you start with the premise that this is largely true during the regular season, one could conjecture that it is less true in a short series as behavior will be highly influenced by the series status--teams down 3-1 behave a lot differently than teams up 3-1 or tied 2-2. This is a classic case of what Bill James called the law of competitive balance.), we have not considered home field advantage, etc. For some more reading on this topic, check out Phil Birnbaum's post at Sabermetric Research and the Baseball Research Journal piece referenced there ("Relative Team Strengths in the World Series" by Alexander E. Cassuto and Franklin Lowenthal, BRJ #35).

Getting back to the actual data, we see what I will call a reverse home field advantage (a 5-game series, in which the "road" team actually hosts 3 games and plays two on the road) 20% of the time, no home field advantage (4 or 6 game series) 41% of the time, and a true home field advantage (7-game series) 40% of the time.

How often does the team with paper home field advantage actually win the Series? Let's break it down by series length:


This is pretty interesting, IMO. The paper home team wins 57% of the series, which seems impressive, but their strongest advantage comes when there is no home field advantage (61%), followed by reverse home fields (56%), and just 53% when there is a true home field.

Of course, the sample sizes aren't great when it's broken down like this, and it is unsurprising that the proportion of series won is less in seven games. What is interesting, though, is that the on-paper home team has such an advantage, and even in series in which they don't really benefit from it in the raw count. Are the first two games at home that much of an advantage, or is there something else going on here?

I'll leave that as a rhetorical question. There are a lot of factors in play here--the sample sizes aren't that large, we have not accounted for the quality of specific teams (which is tough to do in any case because of the fact they play in different leagues which were until recently truly separate in the regular season), etc.--and I don't really want to speculate about the influence of these myriad factors.

I did take a look at the regular season W% of the World Series participants, but as I just said, that's not a particularly telling measure, as it is possible that the leagues were unbalanced in any given year and that a lower W% in one could actually be indicative of a higher-quality team. I checked it anyway, and found that, for the group of series defined throughout this post, the winners had a mean W% of .616 with a median of .616, while the losers had a mean W% of .612 with a median of .610.

Teams with on-paper home field advantage had a mean W% of .615 and a median of .611; teams without on-paper home field advantage had a mean W% of .613 and a median of .610. There's no evidence of any sort of fluky quality difference, at least to the extent that W% captures quality. In terms of W%, the World Series winners, losers, on-paper home teams, and on-paper road teams are all essentially equal.

Let's also break down the series outcomes by on-paper home field advantage coupled with which team had a superior record. These figures will exclude the 1949 and 1958 series as the participants had identical regular season records:


So the team with the worse record has actually triumphed in one more series than their higher W% opponents (for reference, the mean W% for teams with the better record is .635 with a median of .636; the mean W% for teams with the lesser record is .593 with a median of .597, again excluding 1949 and 1958). Teams with home field advantage have been very successful, but those with worse records and home field even more so than teams which had both advantages.

Let's break down the home field W% by each game in the series:


As you can see, games 1, 2, and 6, which are home games for the team with on paper HFA, are the ones with the highest home W%. In game 7, the home field advantage is not particularly large. Those who make a big deal out of WS HFA are fond of pointing out that the home team has won the last eight game 7s, but they were just 2-6 in the previous eight, and I doubt there is anything significant going on. (Although I should point out that the period does correspond to the introduction of the designated hitter in WS play, even if I don't believe that has a significant effect (**)) Between 1952 and 1979 (which includes the 2-6 period mentioned above), road teams were 13-3 in game sevens.

One important caveat on comparing the game-by-game numbers is that as the series extends past the minimum of four games, we should expect to see less of a difference as mismatched teams are eliminated. It doesn't explain why the home field advantages are much smaller in games 3, 4, and 5, though, as there's no reason to suspect that the on-paper road teams are of substantially different quality than the on-paper home teams.

The overall World Series home W% is .573, high compared to the regular season average which is generally somewhere in the neighborhood of .540. Let's use this figure in place of a default assumption of a 50% outcome in each game to model the outcome of a series. Using the combinations detailed above, we can find the probability of any series outcome given these assumptions. For example, the probability of a 4-2 series in which the home team wins games 1, 2, 4, 5, and 6 would be .573^5*.427 (five home wins and one road win). Under these assumptions, we get these probabilities for the possible series outcomes (in this table, "home" refers to the teams with on-paper HFA and "road" to their opponents):


Even using the sample home W% of .573, we only expect the team with HFA to win 52.3% of the time. In fact, teams with HFA have won 56.8% of the series (46 of 79). What is the probability that this could have happened by chance, assuming that 52.3% is the true probability and that each series is independent of the others? It's 12.1%. Even if we assume that there is no true home field advantage at all, and each team will win 50% of the time, there is still a 5.7% chance that 46 out of 79 would be observed.

How about the individual game results (home teams are 268-200, .573)? If the true home field W% was .540 as it generally is for the regular season (and given all the other necessary assumptions for use of the binomial distribution), the probability of 268 successes in 468 trials is 7.1%.

So I am decidedly uncomfortable drawing any conclusions about the strength of home field advantage (on the series or game level) in the World Series from the sample data. The actual results show a stronger home field advantage than we might have expected, but not to such an extent that we must conclude that regular season assumptions about home field advantage do not apply.

It's certainly a good thing to have home field advantage for the World Series, or any game for that matter, and I'm not going to try to argue that basing home field on which league won the All-Star Game is anything but a gimmick. However, given that the previous method of determining home field was simply to alternate it yearly between the leagues, I don't think there's any real harm being done by this approach. If you really wanted to reward the stronger league, the overall interleague record would be far more likely to successfully identify the stronger league, but I don't consider the whole matter worth getting exercised over.

I have posted a Google Spreadsheet with the sequence of games in each series if you are interested. The first group of columns marked G1 through G7 indicate whether the eventual WS champion won the game (W) or lost (L). The second group of columns indicate whether the home team in that particular game won (H) or whether the road team won (R).

Finally, I'll close with some useless trivia. You probably know that there have been three series in which the home team won each game (1987 Twins over Cardinals, 1991 Twins over Braves, and 2001 Diamondbacks over Yankees). The most road games ever won in a series (that I considered for this study) is five, which has happened seven times--1926 Cardinals over Yankees, 1934 Cardinals over Tigers, 1952 Yankees over Dodgers, 1968 Tigers over Cardinals, 1972 A's over Reds, 1979 Pirates over Orioles, and 1996 Yankees over Braves.

P.S. After I wrote this post, but before I published it, Sky Andrecheck published a piece on the importance of World Seires HFA at Baseball Analysts. It addresses an interesting question that I will paraphrase as "Since the Dodgers have such a large lead in the playoff race, is the single most important regular season game left on their schedule (with regards to winning the World Series) the All-Star Game?"

I'll let you read Andrecheck's article to find the answer, but there's one minor point which overlaps with this post worth commenting on. Andrecheck notes that the playoff HFA has been higher than the regular season historically, and reasons that this has to do with the home team being the better team more often than not. While this is true for the league playoffs, there's no reason to suspect it to be true for the World Series in which home field alternated between leagues (even if the All-Star result method of determining home field has the effect of giving on-paper home field to a better team more often than not, home field has not been decided by that rule nearly often enough to have any impact on the results, and the amount of noise involved would be incredible in any event). I don't disagree with the notion that we can't say with any certainty that the World Series HFA is of different magnitude than the regular season HFA, but the better team having more home games leaves a lot to be desired as an explanation (again, for the World Series, not the league playoffs).

He also gives the probability of the team with home field winning as 51.26%, assuming that the home W% in the World Series is 54%. I didn't provide this figure in my post, as I approached the question from the standpoint of "Even if .570 is the true HW%...", but I am in agreement with it (naturally, as it is true by definition given the assumptions we both made).

In the comments to Andrecheck's article, there was a link to Cyril Morong's look at WS HFA, published in 2006, which means that I pretty much repeated here what he had done. However, we disagree on the probability of the on-paper home field team winning the series in six games (and thus of course we also disagree on the probability of them winning the series period). I am pretty sure that this is due to a faulty six-game series sequence he used.

(*) Sorry, I can't help it. I SHOULD take the high ground, but the sniveling "It's not REAL baseball" is way too much for me to handle. I'm weak like that.

(**) There was no DH in the World Series until 1978, at which point it was introduced on an alternating year basis. So in 1978, 1980, 1982, etc. the DH was used in all World Series games, and was not used at all in 1979, 1981, 1983, etc. Starting in 1986, the home team's rules were used.

So while the run of Game Seven home wins begins with the Cardinals in 1982 and also includes the Royals in 1985, in those series the road team's rule was being used in Game 7. All of the game sevens that follow, of course, used the home team's rule.

Tuesday, July 14, 2009

Meanderings

Meanderings are what you get when I either have no coherent ideas for a post or a number of things I want to write about that are all insufficient to fill out a full post. Other times, like this time, it's just a collection of junk thrown together.

* The recent deaths of Ed McMahon, Farrah Fawcett, and Michael Jackson within a few days of each other revived one of my least favorite memes--people dying in threes. I realize that very few people, if anyone, actually takes this sort of thing seriously, and really thinks that if two celebrities die today that movie studios should be contacting their insurance companies. Still, it is a perfect example of how multiple endpoints and loose definitions can lead to some awfully silly things being said.

The endpoints are wide open, as this adage never defines what the period is in which the three deaths should occur. Obviously, if you wait long enough, you will be able to group at least six billion people together in death. Practically, though, it leaves it open until the third person you need to form your group dies. Had Michael Jackson died three days later than he did, he still could have been in the group. If he was still alive and well, then people could have reached back in time for David Carradine, or waited around for Steve McNair and Robert McNamara. No matter.

The loose definition of such groups is also apparent. That they are reasonably well-known is the only qualification. Certainly Michael Jackson's fame outshined the other two, but they are in the group all the same. There was no need to wait around for two other people of Jackson's notoriety. If time had gone by and no one else of note had died, I'm sure somewhat would have dug through the obituaries and found a lesser-known individual to include in the group.

* Speaking of silliness, how about ESPN's 20 Year All-Star team, covering the twenty years that ESPN has been broadcasting MLB games? They have been showing the nominees for various positions during the Monday and/or Wednesday night games, opening up an internet poll throughout the week, and then announcing the winners on Sunday Night Baseball.

Obviously any time you let internet voting occur without any sort of screening or restrictions, you are bound to get some silly results (remember the pitiful All-Century Team that didn't include Hans Wagner among others?) So it's not worth criticizing the selections themselves, and it would be hard to do so anyway because they are the result of a disparate group of individual choices.

However, the whole exercise illustrates why I don't like this kind of exercise when the time period is restricted arbitrarily (obviously ESPN had its reasons for using twenty years, but it has no particular baseball significance). The selection of Nolan Ryan as top right-handed pitcher is illustrative of one of the biggies. Leaving aside the fact that Ryan has been lionized and overrated by many ordinary fans, with his strikeout and no-hit feats overshadowing the more mundane aspects of the game like preventing runs and winning games, and accepting for the sake of argument that Ryan is one of the five or ten greatest pitchers of all time, it is patently absurd to suggest that he is the best right-handed pitcher of the last twenty years, given that he only pitched in four of them...

...Unless you look at it from the perspective of "best to play in this period, period". Since Ryan played in the twenty-year period, he's eligible, and he's a reasonable choice within the bounds of this idiosyncratic viewpoint (remember, above we agreed to accept the premise that Nolan Ryan was one of the very greatest pitchers in history). I don't think this is what most people have in mind when they look at a question like this--do you want to put Cal Ripken or Tony Gwynn on an all-00s team?

There's the middle ground, which would be something like "I'll consider someone if they played a significant amount in the period, whether or not they actually have a case to being the best in that period." From this perspective, you could justify a vote for Cal Ripken on the 20-year team, because he was played in roughly half of the seasons and was still productive in most of them.

And then there's the literalist definition of twenty years, in which only performance within the period is taken into consideration, and thus it is getting dicey when you argue for Nolan Ryan over Dan Haren, let alone Greg Maddux or Mike Mussina. While most people will gravitate towards one of the latter two definitions, these types of exercise usually leave it open-ended, and the results are as much a question of how you approach the exercise as they are a judgment on any of the players involved.

There will be a rash of this stuff coming up near the end of the season and over the winter as the decade ends (Or does it? Even that is not so easy to define). I'll be over here with my fingers in my ear, yelling "STOP!" in vain, thank you very much.

* I love the MLB Network, and think it knocks ESPN's socks off in every aspect of broadcasting, analysis, game coverage, ...except one. Statistics.

The stats displayed on-screen on MLB Network, either during games or on MLB Tonight, are pathetic. I think the standard line for starting pitchers is W-L, ERA, K, and W. That's not so bad except for the omission of innings, which are sorely needed to contextualize the last three categories.

For hitters, though, you get BA, HR, R, and RBI. No plate appearances (or even at-bats). No OBA or SLG. They do display the OBA, SLG, and OPS leaders sometimes on MLB Tonight, but that's about it.

ESPN is running circles around them in this department. The standard batter line when watching a game on ESPN is BA/HR/RBI/OPS, with OBA, SLG, and OPS in tiny print at the top of the screen (at least until the at-bat starts and they are replaced by the always captivating "after x-y count" stats).

* You always see the barb that "you don't watch the games" directed at sabermetricians, and this is often coupled with the "living in your parents' basement" type of stereotype that adds up to nothing more than "sabermetricians are losers". You know, socially maladjusted folks who think girls have cooties and stand in the corner at any sort of social gathering they are roped into attending.

Obviously this argument is not even worth attempting to refute. However, the implicit assumption is kind of funny--that watching a large amount of baseball games makes one cool. After all, this argument is usually advanced by fans, not baseball professionals who are paid to watch and attend games. To the public at large, people who watch a lot of baseball games are probably not considered to be at the top of the social hipness scale. So the whole "watching games" argument (even if one was to accept the premise that sabermetricians don't watch games) really boils down to the Star Trek fans telling the Star Wars fans that they are losers.

* I am embarrassed to say that I was unaware that Steve Phillips attended the University of Michigan. Suddenly, it all makes sense.

Monday, June 29, 2009

Except for the Last Game...

...OSU baseball had a wonderful season, its best since 2003. But the last game was so dreadful that the success of the season will be overlooked by those who don't follow the team closely (which, since we're talking about college baseball, is just about everybody).

The sad truth is that Florida State laid a whooping for the ages on the Buckeyes in the first regional final game, a 37-6 rout in which multiple NCAA Tournament records were broken. It also broke several OSU futility records. OSU has played 3,810 games, and the previous high mark for runs allowed was 24 (incidentally, one of these occurrences was against the Seminoles in 1978, the other against Central Michigan in 1990). The worst loss in OSU annals was 20 (23-3 to Miami-FL in 1999 and 22-2 to New Mexico in 2002).

There's no doubt that the last act was a downer, and it highlighted the team's one glaring weakness (pitching depth), which had been on display all season long. However, the carnage should not be allowed to overshadow the fact that OSU was one of the last 32 teams standing in the NCAA Tournament and had dispatched Georgia earlier in the day to finish second in the Tallahassee regional.

OSU started with a fairly soft non-conference schedule, but it set the team up well for conference success. The Bucks won their first seven contests and were sitting at 18-3 entering Big Ten play. The biggest non-conference game was at Miami-FL, where Alex Wimmers pitched five strong innings on three days rest in a 7-1 victory.

The Big Ten switched from a four-game weekend series (with a Saturday doubleheader made up of seven inning games) to a standard three-game series. This change was fortuitous to OSU in this particular season as it reduced the importance of second-line pitchers. The Bucks won a series at Penn State and lost at Minnesota to open the conference season. At 3-3, they were in the middle of the early pack and lost ground to Minnesota, who figured to be one of the top contenders.

The Bucks were able to jump into first place with back-to-back sweeps (MSU at home and Purdue on the road), but even after taking two of three in the next three series (Northwestern, the forces of evil, and at fellow contender Illinois), they had fallen out of first by a half-game to Minnesota.

In the final weekend, OSU was able to sweep Iowa at home, meaning they would need just one Penn State win over Minnesota to take the title. On Sunday, Penn State obliged, and the Buckeyes had their first regular season Big Ten title since 2001.

Meanwhile, the weekday games (generally non-conference home games against weaker area opponents) had not been going as smoothly as usual. OSU's record in these games was just 3-6 (although two of the losses came at Louisville, a strong club and an exception to the normal scheduling).

Here is Ohio's record in non-conference home games (essentially mid-week games, with exceptions like Louisville) over the last five seasons:

2005: 6-1
2006: 5-2
2007: 5-4
2008: 6-3
2009: 3-4

This is another manifestation of the pitching problems (of course, the sample size of seven games could be a factor as well). Winning the Big Ten title with a sub-.500 record in those games is an oddity.

The Big Ten tournament was held in Columbus; under the old format, which was in place until this season, it would have been at Bill Davis Stadium. But for the first time the Big Ten tried a neutral site, off-campus location for the tournament, namely the Clippers' brand-new Huntington Park. OSU beat Illinois in the opener 7-4, but Indiana and Minnesota each roughed up the Buckeye staff (13-3 and 9-6) to consign them to a third-place finish.

In the NCAA Tournament, the Bucks were made a #3 seed in the Tallahassee region with #1 Florida State, #2 Georgia, and #4 Marist. A disastrous eight run first inning enabled Georgia to cruise to a 24-8 win (again, the second-line pitching was brought in to be drubbed), but OSU stayed alive with a 6-4 win over Marist. Meeting Georgia for the right to play for the regional title against Florida State, the Bucks rallied from a 4-0 deficit for a 13-6 win. The season ended with the FSU game, already covered above.

The Buckeyes finished first in the B10 with a .689 EW%, fifth with a .540 EW% (Minnesota led at .635), and fourth in PW% (Indiana led at .639). Of course, these measures were affected by the blowout loss to FSU and some other blowouts throughout the season (not all of which went against Ohio, of course, but more did than didn't).

As a whole, the defense allowed 7.3 runs/game, ninth in the Big Ten--this despite the fact that Bill Davis Stadium is a fairly strong pitcher's park. The pitching staff had two shining bright spots. The first was sophomore Alex Wimmers, who went from a potentially brilliant but maddeningly inconsistent middle reliever to staff ace. Wimmers had multiple double digit strikeout games, pitched the first nine-inning no-hitter in OSU history on May 2 against the servants of evil, was the co-Big Ten Pitcher of the Year, and was named a first-team All-American by PING. He averaged 11.7 K/9 and was 35 runs better than an average Big Ten pitcher.

Senior closer Jake Hale also had an amazing season. A starter as a freshman and junior and the closer as a soph, he returned to the pen and set the OSU single season saves record with 18, appearance record with 40, and career save record with 29. He had a 11.0 K/9 and was +25 RAA in just 55 innings, and was drafted by Arizona in the 27th round.

Beyond that pair, OSU had three pitchers of moderate effectiveness. Sophomore reliever Drew Rucinski led the team in wins at 12-2, and early in the year was lights out. He faded a bit down the stretch, and ended at just +3 RAA (Rucinski's season stats, like those for most of the pitchers that follow, were hurt by their battering in the Florida State game, which for many of them was their third appearance in a must-win game in two days.)

The other two starters, sophomore Dean Wolosiansky and junior Eric Best each checked in at -3 RAA. Past that, the pitching was a disaster, as six other pitchers combined for 133 innings and a 13.53 RA. Most disappointing was that sophomore lefty Andrew Armstrong (penciled in the rotation with Wimmers and Wolosiansky) was derailed by injuries as was strong-armed freshman righty Ross Oltorik (a walk-on quarterback in the fall).

Coming into the season, I and other followers of the team felt that pitching depth might be a strength. Oops. However, most people felt that the offense would continue to be mediocre and woefully lacking in power, and that was decidedly not the case. The Bucks scored 7.9 runs/game to pace the conference, led in BA at .328, were second in OBA at .390 (Purdue, .395), and first in SLG at .495. Rather than lacking in power, the Buckeyes were second in the conference with 118 doubles (Indiana, 119), first with 66 longballs, and first with a .168 ISO.

Sophomore catcher Dan Burkhart had a breakout season, taking Big Ten Player of the Year honors with a .354/.438/.589, +21 RAA performance and solid defense behind the plate. At first base, sophomore Matt Streng was forced into action and acquitted himself fairly well with a surprising 8 homer, +3 RAA campaign. Junior second baseman Cory Kovanda did a great job getting on base, tying for the team-lead with 33 walks, finishing second to Burkhart in OBA at .427, and winding up +10 RAA despite lacking power (.101 ISO). At shortstop, sophomore Tyler Engle didn't do much well except get on base, but drawing 28 walks in just 130 at bats enabled him to post a respectable .285/.411/.423, +3 season.

Senior captain Justin Miller was ice cold early in the season, and wound up at third base after the expected starter, junior Brian DeLucia, went down with a finger injury. Miller closed his career with a .310/.369/.506, +4 season. The key infield reserve was junior Cory Rupert, who saw significant time at short and third, struggling to .279/.329/.388, -8.

Junior left fielder and leadoff man Zach Hurley teamed with Kovanda to form a dynamic on base duo at the top of the lineup (.346/.421/.510, +15), and was picked by Florida in the 45th round of the draft. Junior college transfer Michael Stephens (a junior in eligibility) manned center ably and while billed as having doubles power, paced the team with 14 round-trippers and wound up at .346/.375/.608, +13. The hole in his game was his walk rate, with just 11 in 237 at bats. In right field, senior Michael Arp was not particular productive, hitting .295/.345/.405, -7 RAA. Junior DH Ryan Dew finally delivered on his offensive promise, hitting .388/.420/.562, +16. Those four combined to play almost all of the available innings in the outfield.

The early outlook for the 2010 Buckeyes is bright. Most of the key players return, with Jake Hale, Justin Miller, and Michael Arp the only senior starters. Given Hurley's low draft position, it is likely that he'll be back as well. Hopefully OSU will be able to find some freshman pitching that can at least soak up innings, and get Armstrong and Oltorik healthy to go along with Wimmers, Wolosiansky, and Best. This season showed that it's possible to achieve great things with a paper thin staff, but it's not a feat that Bob Todd will want to attempt again any time soon.

Wednesday, June 24, 2009

I Thought There Were Nine Players on a Team...


Alternate title: The Indians lost to a team with a 7-man lineup? Figures...

Monday, June 22, 2009

And Again...

Today I received my copy of the Baseball Research Journal, 2009 Volume 1 from SABR. As I have said before, the increase in the quality of the BRJ from the mid-90s to now has been remarkable, and much credit is due to the contributors as well as to SABR's publication directors. I look forward to reading it and I may devote a post down the road to my comments on particular articles.

However, escaping the continuing reuse of a particular offensive statistic appears to be too much to ask for. Quickly skimming the statistical pieces, I came across "Offensive Strategy and Efficiency in the United States and Dominican Republic" by Robert J. Reynolds and Steven M. Day, which appears to be a study of the shape of offensive performance of American and Dominican batters. Just flipping through it, the graph "Plate Appearance Base Average" caught my eye and raised my suspicions.

Sure enough, my suspicions were confirmed. This is yet another presentation of bases/something, this time bases/plate appearance. The authors write "...we use on-base percentage and introduce the statistic plate-appearance base average as a plate-appearance analog to slugging percentage. PABA is calculated as the sum of the bases achieved in three categories--hitting (TB), BB, and HBP--divided by the total number of plate appearances: (TB + BB + HBP)/TPA...PABA is similar to bases per plate appearance and runs created, though these later include stolen bases, and advancing other players through sacrifices."

At least one can rejoice that they are not claiming to be the originators of this statistic as many others have. It appears as if they use both OBP and PABA to determine overall effectiveness, which at least somewhat defuses a major downfall of bases/PA in isolation, which is that it does not properly account for outs. Still, it would be nice to be spared the "introduction" of bases/PA, and it would be nice to see a better metric used as the primary offensive measure.

Last week, I linked to a new website, the Barry Code, which features the work of Barry Codell, the creator of Base-Out Percentage, which was one of the first bases/something metrics and the first bases/out metric. Many others have come along with bases/out metrics (most famously Tom Boswell's Total Average), but Codell was the first. While I have my qualms about any base/something metric, bases/out is certainly the "better" form (better in the sense that it better captures a player or team's true offensive efficiency), and at the very least was a nifty idea at the time of its introduction. It is a continuing frustration of mine that people keep re-"inventing" these measures, with seemingly no knowledge of the work of others that went before them, which in the internet age can be discovered with a cursory Google search.

To be fair, Messrs. Reynolds and Day don't fall into this category exactly, as they make no claim to developer status. Still, the relentless recycling of bases/PA (and re-invention of bases/out) remains one of my top sabermetric pet peeves.

Wednesday, June 17, 2009

I guess it's like the old College All-Star Game now, huh?


Pointed out by "Roid Vulture" at BTF.

Monday, June 15, 2009

Mid-Season Managerial Changes, 1982-2008

Disclaimer: This is not so much a study as it is a collection of data. There is no claim that the data is statistically significant; although I will use it in the course of discussion, I am not making any formal claims. You will also note that I have included a number of graphs; they don't do much for me (I'd rather just have the data table), but some readers may find them helpful in this case. With any of the images, you can click on them to enlarge as they may be tough to read otherwise.

I started with 1982 because it seemed like a good cutoff point--I didn't want to go too far back, and strike years cause a bit of a problem. I'm certainly not claiming that there is any fundamental difference with regard to managerial dismissals between, say, 1978 and 1982.

I counted all permanent managerial changes that occurred during the season with four exceptions, identified by either the Sports Encyclopedia: Baseball or my memory as not baseball related. The three exceptions are:

1. Dick Howser, KC 1986--medical issue
2. Pete Rose, CIN 1989--banned from baseball
3. Tommy Lasorda, LA 1996--medical issue
4. Larry Dierker, HOU 1999--medical issue

Only the first change is counted for any team-season. If there is an initial interim replacement, and later a permanent replacement, I have lumped them together as most of the interim stints are just a couple of games. I have tried to use the word "change" primarily, but sometimes I have lapsed into "fired", even though some certainly were resignations (and unless my memory from two years ago has been completely fried, Mike Hargrove's departure from Seattle really was a resignation). I'm not using "fired" as a technical term, here, okay?

I have a link to the spreadsheet I used at the end of the post, if you are interested. I am going to do this in a Q-and-A format:

At what point in the season did managerial changes occur?

The earliest (in terms of games) changes came after six games: Cal Ripken (BAL, 1988) and Phil Garner (DET, 2002). Each team started 0-6.

The latest change came after 160 games, when Larry Bowa (PHI, 2004) was let go and Gary Varsho managed the final two games.

The average change came after 80 games, which seems logical. The median was 75.5 games. No team made a change at the exact halfway point of 81 games; in 1990 both Jack McKeon (SD) and Whitey Herzog (STL) were replaced after 80 games, while Bob Boone lasted 82 games for KC in 1997 and Jerry Narron the same for Cincinnati a decade later.

Here is a table showing the number of games elapsed when a change was made. "0" means that the change occurred after 1-9 games; "10" after 10-19; and so on:


And a graph of the same:


The pattern seems to be a lull after the All-Star break; if you make it to the halfway point, are relatively safe for a month, month and a half. Then things pick up again towards the tail end of the season.

Which franchises made the most changes?

During the period in question, every major league franchise has made at least one mid-season change. I expected that the team with the most would be the Yankees, but I was wrong--they are in an eight-way tie for fourth with five changes. The Reds have changed managers seven times mid-stream since 1982 (eight if you count Rose's banishment):

1982: Russ Nixon replaced John McNamara
1984: Pete Rose replaced Vern Rapp
1993: Davey Johnson replaced Tony Perez
1997: Jack McKeon replaced Ray Knight
2003: Ray Knight and Dave Miley replaced Bob Boone
2005: Jerry Narron replaced Dave Miley
2007: Pete Mackanin replaced Jerry Narron

While there are five teams that made just one mid-season change, but three are fourth-wave expansion teams (and two of them have pulled the plug on their manager in 2009). The two longstanding franchises that made just one change are Pittsburgh (Pete Mackanin for Llloyd McClendon, 2005) and Los Angeles (N) (Glenn Hoffman for Bill Russell, 1998).


Has the frequency of mid-season firings changed over time?

Indeed it has, and the change to the division/playoff format of 1994 *appears* to be a reasonable explanation for the altered behavior. Here are the changes by year; N is the number of major league teams:


The maximum of 31% (8 of 26) was reached in both 1988 and 1991. The only seasons without any changes were 2000 and 2006. Here is the percentage in graph form:


I also figured a moving three-year average and produced a graph (the years on the x-axis are the first years of the three-year period):


1992-1994 saw a nosedive in the frequency of firings, one from which there has been a bit of a recovery, but never to a frequency any higher than the 1991-1993 period. 1994 certainly poses a problem due to the strike (and 1995 to a lesser extent with a 144 game schedule), but even if you ignore the three-year periods starting between 1992 and 1994 (i.e. those that include 1994), the rate of changes has dropped.

It certainly seems logical to me that the existence of four additional playoff spots led to a reduction in firings. More teams remain in the hunt despite slow starts, and a slow start is easier to overcome.

Summing it up, in the period 1982-1993, 19% of teams fired their manager mid-season. From 1996-2008, that rate has fallen to 10%. A difference of about 9% in a league of thirty teams is three (2.7) fewer changes per season.

Did teams improve their record after the change?

Yes, they did. The composite record prior to changes was 3562-4580 (.437); after changes it was 3881-4390 (.469). The total season record for the teams was 7443-8970 (.453).

76 of the 102 teams had a better record after the change than before (75%).

Of course you have to be very careful with this data. Cito Gaston was fired with a 72-85 record in 1997 and Mel Queen took over and went 4-1. Thus the 1997 Blue Jays count as a team that improved their record, but obviously one would not want to draw any conclusions from five games. Teams like that also can cause the aggregate records to be distorted.

Nonetheless, I am comfortable with the conclusion that teams generally had better records post-change. The improvement from aggregate wins and losses was .032. The average improvement (weighting all teams equally, even teams like the Gaston/Queen Jays) was .055. The median of the same was .045.

A better approach might be to take a weighted average, with the weight determined by the minimum of games before/after the change. Gaston managed 157 games and Queen managed 5, so the 1997 Jays will be weighted at 5. A team in which the change was made at the exact halfway point would get the maximum possible weight, 81.

Doing it this way, the average improvement is .046. Attempting something else in lieu of more advanced mathematical techniques, one could try weighting by the harmonic mean of games before and games after (2*before*after/(before + after), which you may recognize as Bill James' Power/Speed Number. It's also 2/(1/before + 1/after)). Done in this manner, the weighted average improvement is .049.

But don't teams that fire their managers generally feel as if they are underperforming? Could some (most? all?) of the difference in performance after the change be a result of regression to the mean?

Now that is a good question. And yes, I believe that is what is really going on here, although what follows in no way proves it.

What I would really like to do, if I had a lot more patience for this kind of thing than I actually possess and a better database, is this:

1) figure an expected record for each team in MLB during this period
2) for each team that made a managerial change, find a team (or teams) with similar records and expected records at the point at which the change is made
3) compare the performance of the teams that made a change with those that stayed the course

I *suspect* that if one did such a study, they would find that the performance of the two groups of teams was very similar, and that there was little proof that the managerial change was the impetus for improvement.

I do have a poor man's study here for you, though. In the Bill James Guide to Managers, James figured an expected record for each team based 50% on the previous year's record, 25% on .500, and 12.5% on each the second and third most-recent records. For example, the Angels played .438 ball in 1993, .409 in 1994, and .538 in 1995. So their expected record for 1996 was:

.5(.538) + .25(.500) + .125(.438 + .409) = .500

I figured expected records in this manner for each team that made a managerial change. Obviously, this is a crude approach and thus the study built on it is crude.

The teams that made managerial changes had a combined expected record of .492. Before the change they were had an aggregate W% of .437; after, .469; and for the season as a whole, .453. At the time of the change, 91 of the 102 teams had a lower than expected record (89%).

So one might well expect that many of these teams would improve on their own, whether a managerial change was made or not. It is of course impossible to say to what extent that is true . One must grant the possibility, however far-fetched it may be, that these managers were all an albatross around the neck of the club, dragging it down and preventing it from reaching its true potential. I don't buy it, certainly not in the majority of cases. Managers are relatively fungible, and so they are offered up as penance for a poor season, demonstrating to the fans or the players or the media that the brass is being proactive.

The teams went from playing at 89% of expectation before the change to 95% after. Even after the change, the teams did not play up to expectations, but the method of setting expectations is nowhere near accurate enough to get carried away with this tidbit. Ideally, the actual rosters would be used to set expectations, and you would account for injuries and the like.

Here is a spreadsheet listing all of the changes, along with the expected records.