Thursday, January 31, 2008

Sabermetrics Wiki

Tango Tiger has set up a sabermetrics wiki here. Hopefully, this will become a one-stop sabermetric encyclopedia and research diffusion center. For more information, see this thread, and then come over and contribute.

Tuesday, January 29, 2008

Best Hitting Infields and Outfields, 2007

In a BTF thread on the Phillies’ signing of Pedro Feliz, there was some discussion of whether or not he along with Howard, Utley, and Rollins gave them the best hitting infield in baseball. Obviously, that is a foreword-thinking question, but I thought that since I had the data by position for each team readily available, it would be mildly interesting to see which units performed the best last season.

The stats below are from the Baseball Direct Scoreboard; infielders are (duh) first baseman, second baseman, third baseman, and shortstops. Pitchers, catchers, and DHs do not factor in to the totals for either the infield or the outfield. All players who were classified by STATS as playing the position are included in the totals. There are no park adjustments, and RC does not include basestealing.

Beginning with infielders, three teams were within one-tenth of a run per game of the lead: the Yankees (6.19 RG), the Marlins (6.18), and the Phillies (6.14). Since the comments about Feliz inspired this little survey, it is worthwhile to note that third base was definitely the weakest infield position for Philadelphia last year (4.01)--however, Feliz himself created just 4.07 runs last year.

On the other side, the Giants trailed (3.90), followed by the Twins (4.00) and the White Sox (4.04). The overall average was 5.03.

The top hitting outfield was that of the Rockies (6.63), well ahead of the Tigers (6.34) and the Phillies (6.23). However, with a park adjustment, the Rockies drop to 6.03.

The average outfield created 5.17 runs. The worst hitters were on the south side of Chicago, as White Sox outfielders limped in at 4.34, ahead of the Mets, Diamondbacks, and Royals (4.40).

Monday, January 28, 2008

Beating a Dead Horse, pt. 1

Last summer, I had a post that I described as a rant about OPS, and my problems with the stat. I got (relatively) a lot of feedback about that piece, and you can add that fact to other evidence that shows OPS is still a topic of interest among people who are interested in sabermetrics. If you go to the kinds of discussion forums in which baseball fans of a sabermetric bent, but not themselves hard-core sabermetricians post, various combinations of OBA and SLG, along with rate stats revolving around bases gained, remain a popular topic.

More than one of the respondents to my post felt that it was in essence beating a dead horse, because all of the serious sabermetricians of the world have moved on and found more important things to worry about, and that enough has already been written on the topic to satisfy the inquiries of the less-informed. I largely agree with this; however, I am fairly stubborn by nature, and I like to express myself on sabermetric topics myself, even if others have already covered the matter thoroughly. So despite the fact that I agree that OPS is a topic whose time has passed, I am going to write a little series on it and its cousins. Also, I feel that if people are going to use OPS, they might as well understand how it relates to run scoring rates.

I would also add that by posting about it here, I do not expect anyone to pay attention to it. If you feel that horse has been bloodied enough, then by all means, pay no attention to this post. If no one else ever reads this, that would be just fine with me.

The earlier linked post did not deal a lot with numbers, which these will have more of. I don’t want to rewrite the same thing over again, but I do want to summarize my main points:

1) OPS is a quick and dirty statistic, and it is alright for this purpose. I am not saying that you should never, ever look at OPS.

2) OPS is a statistic which is unitless. If you write the formula over a common denominator, you have ((H + W)*AB + TB*(AB + W))/(AB*(AB + W)), which is a bunch of gobbledy-gook.

3) OPS also does not come in an estimated unit. Runs Created is an example of a stat that does; (H+W)*TB/(AB+W) is an estimate of runs scored. Obviously the result is not actually runs scored, thus it is an estimated unit (see my previous post for more on this). OPS does not have an estimated unit. “The higher the OPS, the better” or “OPS is an estimate of a player’s offensive productivity” are both true statements, but they do not confer a unit upon the stat.

4) OBA is a fundamental baseball statistic. If OBA did not already exist, you would want to invent it. SLG is not a fundamental baseball statistic, because it does not measure any fundamental quantity--it measures “bases gained by the batter on hits”. This is a unit, and it is a useful thing to know, but it is not fundamental. If this does not convince you, ask yourself the question, “What does SLG represent?” Some people will be tempted to say “power”, but that’s obviously false since it includes singles. Some people will say “advancement of baserunners”, which can be true, but it is also true that it is not even close the being the most accurate estimate of advancement. In contrast, there is no definition for a statistic that would better define what OBA attempts to define than OBA itself, at least that can be derived from the official statistics.

5) Because of points 2, 3, and 4, there is nothing special about OPS. When Pete Palmer comes around and invents “OPS+”, it may make sense to claim that the name is a little misleading, but it doesn’t make any sense to complain that it causes distortions in measuring performance because it deviates from OPS. Of course, it makes even less sense to complain about it when it was Palmer who invented both OPS and OPS+.

With that out of the way, let’s talk math. Throughout this series, I have defined OBA as (H + W)/(AB + W); I did not mess with HB or SF, so keep that in mind. Also, I will be focusing on things on the team level, and running a lot of regressions. Then I will be testing the accuracy of the equations on the same sample from which I derived them. I recognize that this is not the best approach to take, but I think that if you focus on the relative accuracy of the formulas to each other rather than the absolute RMSE figures, you will not be mislead too far.

Let’s start with the premise that we have OBA and SLG data for a large group of teams, and we want to estimate a run scoring rate from them. In this case, our large sample will be all teams 1961-2002, except 1981 and 1994. We will also look at using relative OBA, SLG, or OPS to predict relative runs scored, as established by the composite average for the dataset. I am doing it that way because that is how OPS+ is expressed, and because using a constant league average will wash out the adjustments.

Let’s define aOBA as OBA/LgOBA, aSLG as SLG/LgSLG, aOPS as OPS/LgOPS (this is what I called “SOPS+” in my earlier post), aR/P as (R/PA)/Lg(R/PA), and aR/O as (R/O)/Lg(R/O). OPS+ is OBA/LgOBA + SLG/LgSLG - 1, which is the same as aOBA + aSLG - 1.

When we regress OPS to estimate runs, what run rate should we regress to, R/PA or R/O? We can all agree that the most important thing to know about a team’s offense is its R/O, so that would seem to be the right choice. While this does not transfer perfectly to individuals, it is still true that R/O is more telling for them than R/PA is, and R/O is generally fine to use as an individual rate stat.

So let’s look at the equations to predict aR/O from aOPS and OPS+ for the sample in question:

aR/O = 2.06(aOPS) - 1.06
aR/O = 1.06(OPS+) - .06

Here we can see that OPS has a 2:1 relationship to runs scored. If you are 5% better than the league average in OPS, you will be approximately 10% better in runs scored per out (and, by extension, in runs scored). On the other hand, OPS+ has an almost 1:1 relationship to runs scored.

The practical implications of this are that if you see a player listed with an OPS+ of 125, you can interpret this as “the player is estimated to create 25% more runs/out than the league average.” It is of course an estimate, and it may not be as accurate as other estimates, but it does scale properly.

You cannot do the same thing with aOPS (and since aOPS is simply OPS divided by a constant, the same goes for OPS as well). If you have a batter with an OPS of 900 in a league with an OPS of 750, saying that his aOPS is 120 means nothing other than that his OPS is 20% higher than the league average. It does not mean that he created 20% more runs--in fact, he created something close to 40% more runs.

As I mentioned in the older piece, people are conditioned to expect that when they see a stat called X+, it will be calculated as X/LgX. OPS+ breaks this convention, and it is curious why Pete Palmer chose to name it as he did (originally it was PRO+, but OPS itself was PRO, so it was the same situation). However, Pete’s choice to use OPS+ instead of aOPS was a good one, as it is more accurate and expressed in estimated units that have meaning.

Once we have the above equations, we can estimate team runs scored and see how accurate the estimates are (keeping in mind the caveats about using the regression on the sample it was derived from). We can find Runs = aR/O*Lg(R/O)*O. We know that for our sample, BA = .258, OBA = .324, SLG = .391, R/PA = .117, and R/O = .172. Plugging everything in, the RMSE against actual runs scored is 25.72 for the aOPS equation and 24.88 for the OPS+ equation.

What would happen if we tried to predict R/PA, rather than R/O? We would get these equations:

aR/P = 1.74(aOPS) - .74
aR/P = .90(OPS+) + .10

The estimate for team runs scored will be Runs = aR/P*Lg(R/PA)*PA. The RMSEs in this case are 23.76 for the aOPS equation and 23.64 for the OPS+ equation--over a run better than the R/O predictions. Why is this?

First, let me claim without presenting any evidence that most statistics do better at predicting team runs scored when estimating R/PA than it does when estimating R/O. To understand why this is, we need to remind ourselves about the relationship between R/PA and R/O. Assuming, as we have in this case, that the only outs are batting outs and that there are no ways to reach base that are not included in OBA, the relationship R/O = (R/PA)/(1 - OBA) holds. This is not an “estimate”; it is a demonstrable mathematical truth. As you can see, the On Base Average is key, since it is the complement of the rate at which outs are made. It is better to estimate R/PA from OBA and SLG, then convert it to R/O by dividing by (1 - OBA). Instead of doing a regression to try to incorporate the value of OBA, you are better off to use OBA directly for that purpose.

So it is in fact more accurate to estimate R/PA from OPS than it is to estimate R/O. However, if you agree with the premise that individual productivity should be measured in terms of runs/out, and you use OPS or OPS+ as your rate stat of choice, you are in essence locking yourself in to considering the less accurate R/O relationships.

Sunday, January 20, 2008

Units and Comparisons

In writing some other stuff, I’ve noticed that I’ve been referencing this topic, and so I figured I should just write about it in isolation so that I don’t have to go off on a tangent elsewhere.

This kind of goes back to Bill James’ article in the 1987 Baseball Abstract about “meaningful and meaningless statistics”. When looking at a baseball statistic (which I’m using to mean a category like “Triples”, not a single statistic like “Curtis Granderson had 21 triples in 2007”), I ask myself a few questions. These questions don’t answer how worthwhile it is to know, but they are very useful in considering derived statistics:

1) What are the units of the statistic? Are they actual units or estimated units?

Consider a fairly mundane counting statistic, the balk. The unit of a balk is clearly “balks”, and since the category is simply a count, they are actual units.

Batting average is measured in units of “hits per time at bat”. Again, they are actual units, although the at bat itself is a weird, kind of artificial subcategory of plate appearances.

OPS is measured in units of “total bases per at bat plus times on base per plate appearance”. Since we have two different denominators, there’s not a clear unit here as there is in the two components taken separately. The unit total bases per at bat plus times on base per plate appearance has no clear meaning; people use the stat as an approximation of overall offensive ability, but overall offensive ability is not a unit either. OPS has no units.

Runs created is measured in units of “estimated runs”. While both RC and OPS are estimates, the distinction between these will be explained below.

Questions two and three apply primarily to derived statistics, not from counts of events.

2) If the derived statistic is measured in an estimated unit, is it one that is fundamental to our understanding of baseball?

By fundamental, I mean units of things that really matter in terms of winning baseball games. If the stat is expressed in estimated runs or wins, then it is fundamental, although I would consider other things to be fundamental.

For example, On Base Average is a very fundamental thing to know; the rate of reaching base. Of course, OBA is not really an estimated unit, but an actual count of things.

I also consider any sort of event frequency with a sensible denominator to be fundamental...walks/PA, homers/PA, hits/balls in play, etc. Of course, some of these are more telling than others (catcher’s interference/PA is not particularly important to know), but they all are very straightforward, basic pieces of information.

Slugging Average is a more interesting case; clearly “bases gained by the batter on hits” is a factual count. However, bases gained by the batter on hits is not critical to understanding baseball, like the rate of reaching base is. If it were “bases gained on hits by batter and baserunners”, then it would be a bit more telling, either on the team level or as an estimated unit on the individual level. So I’ll leave that one up there as one to be decided on. It’s sort of like if you took (balks + wild pitches)/inning.

OPS fails this test, since it’s not measured in terms of anything. That does not mean that OPS cannot be transformed by mathematical operation into a fundamental estimated unit, like runs, but on its own, the units are meaningless.

Win Shares is measured in wins, except the wins are multiplied by three. I consider that to be fundamental. If you want to be a stickler and demand that Win Shares divided by three to consider it a fundamental estimated unit, that’s okay too. The reason I don’t draw a distinction is that the transformation is scalar and straightforward.

3) Can two players (or teams, etc.) be compared using this statistic by the difference between their figures? By the ratio? Both? Neither?

What I am getting at here is does the difference or the ratio have meaning, other than just to tell us which is better. For instance, an OPS of 1000 can be compared to an OPS of 800, and seeing that the one player is +200 points or has a 1.2 ratio, we can see that the 1000 OPS is superior. But the +200 points or the 1.2 ratio don’t have any meaning other than facilitating the comparison.

For an example of a statistic that can be compared by difference and ratio, take runs created per game. 5 RG is two more runs per game than 3 RG, or it’s 67% more runs. Either way, the numeric result of the comparison is expressed in a meaningful unit. There are many stats that would fall into this category: Winning Percentage, On Base Average, Wins, Losses, …, Balks.

However, there are some derived statistics for which only one of the operations produces a meaningful result. Take Runs Above Average for example. If I tell you that one player is +10 RAA and the another is +1, then we have a ratio of 10 and a difference of 9. The difference of nine tells us that Player A contributed nine more runs beyond an average player than Player B did, which is valuable. But the ratio of 10 just obfuscates things, unless you take the position that RAA measures value, and thus Player A is ten times more valuable. Even if one takes that view, the ratio gives a much cloudier picture of the disparity in value than does the difference.

Statistics with meaningful ratios but meaningless differences are harder to come by, but one example is ERA+. ERA+ inverts the usual format of a relative statistic (X/LgX) to LgX/X, in order to make a figure above 100 desirable as it is for OPS+, or Relative Batting Average, or any number of other such stats.

This seems innocuous enough at first glance, but it causes some problems, and you need to be careful when averaging it, as Tango Tiger has shown. Suppose that you have a pitcher who works exactly 200 innings in consecutive seasons in leagues with an ERA of 4.50. In the first season, our pitcher’s ERA is 3.00, and thus his ERA+ is 1.50. In the second season, his ERA jumps to 4.00, and his ERA+ comes in at 1.125. Since he worked the same number of innings each year, we can just average his ERAs together and find that he has compiled a 3.50, which is a 1.286 ERA+. However, if we average his ERA+s, we get 1.313.

The reason this ends up happening is that when you invert the calculation, earned runs, rather than innings, become the denominator quantity. So in order to average the two seasons, you must weight by earned runs (and indeed (4*1.125 + 3*1.5)/(4+3) = 1.286).

For the same reason, the difference between the two means nothing. 1.5-1.125 = .375. What does .375 represent? If we take it times the league ERA of 4.5, we would expect to get the difference between the two ERAs (1), but instead you get 1.69.

Suppose that we instead consider ERA/LgERA. The seasonal figures are now .667 and .889, and we know the total to be 3.5/4.5 = .778, and the average of .667 and .889 is indeed .778. Now the difference between the two, multiplied by the common league ERA is (.889-.667)*4.5 = 1.

The ratio between the two is .889/.667 = 1.333; the ratio between the ERA+s is 1.50/1.125 = 1.333. You can see that the ERA+ ratio is meaningful, but the difference is not, and that’s something you should always keep in mind when working with ERA+ over the course of a pitcher’s career. So while we may have become accustomed to ERA+, its reciprocal is easier to work with and is meaningful in ratio and differential comparisons.

Then there are the statistics for which neither the difference nor the ratio has any intrinsic meaning. Generally speaking, these are the derived stats that I cannot stand, and wish that their inventors and figurers would convert them to a different format. Examples include OPS (since it is unitless to begin with), EQA, and Offensive Winning Percentage. I’ll examine the case of OW% here. OW% starts with a statistic, Adjusted RG, which is meaningful by both difference and ratio, and converts it into a format where it is meaningless by both, at least for application to individual players.

Assuming a pythagorean exponent of 2, OW% = RG^2/(RG^2 + Lg(R/G)^2), or ARG^2/(ARG^2 + 1). Suppose we have a hitter with an ARG of 110, and another with an ARG of 120. Player A has created 110% of the league average runs per out; Player B 120%. The ratio 120/110 is meaningful; it tells us that Player B created 9% more runs per out than did Player A. The difference 1.2-1.1 is meaningful as well. If we multiply by the league average R/G (say .18), we will find that player B was created .018 runs/out more than Player A.

When we convert to OW%, Player A is now at .548 and Player B is at .590. The ratio between them is now .590/.548 = 1.078. In what sense was Player B 7.8% more productive, more valuable, more whatever than Player A? The answer is in no sense that reflects on their actual status as individual members of a ballclub. It is true that a whole team that hit like Player B would be expected to win 7.8% more games than Player A. However, this does not translate directly into any statement of their actual value as individual members of a team.

The difference is just an unintelligible. OW% is okay for the thought exercise aspect of “how good would a whole team of this guy be”, but when it comes to actually measuring the value of a player to his team in a meaningful way, it fails. All of this is not to say that non-linear relationships have no place in evaluating players, but if you’re going to use them, you need to be sure that you model reality and not an unrealistic scenario. For example, you could ask “What would the team’s W% be if it was made of eight average players and Player X” and use the pythagorean formula to make an estimate. The relationship between two player’s OW%s figured that way would not be the same as the relationship between their ARG, but one could argue that it would be a truer reflection of their value. OW% cannot make that claim, and thus the non-linearity just serves to obfuscate relationships between players.

To wrap this rambling atrocity up, the most useful statistics tend to have positive answers to all three questions: they are denominated in some sort of unit, that unit is fundamentally important, and the ratio or difference between players or teams express the relationship between them in a meaningful way.

Monday, December 17, 2007

Providing Zero Insight, but Filling Space Nonetheless

In Bill James’ early Abstracts, he had a little box entitled “Talent Analysis” for each team that estimated the composite value of all of its players (as estimated by the Approximate Value method), what percentage of it was acquired through various means (trades, free agency, development), what percentage of it fell into defined age categories (“young”, “prime”, “past-prime”, and “old”), and how much total value had been produced by the team’s farm system. This article is going to be in the spirit of those and look solely at the offensive players for 2007, and without regard to ho w players were acquired, only whether they were products of the farm system or not. Also, I have omitted the age breakdowns, although I may look at that in the future.

I should note that I do not consider my examination here to be particularly insightful, and certainly it is not unique. This is the kind of stuff that I sometimes figure myself and keep to myself, but since I still can’t (or, more accurately, want to go through the effort to do so) access my pre-written articles, I have to fill this space up with something.

I also would be remiss if I did not point out that this strain of analysis did not die with the Abstract, but is in fact being practiced in other places, most notably by Steve Treder at the Hardball Times. So not only am I just ripping off James’ idea, I’m covering old ground.

Now, for a long list of caveats. One is that I have only considered hitters, and furthermore only hitters with 300 PA. So there are guys who were injured or who were part-time players or what have you and really did have value that are being ignored here.

Second is that I have used my own personal WAR figures, except I have multiplied them all by seven, and put a floor of zero on them. I did this because I didn’t want to get caught up in the numbers as WAR, but wanted something with a direct linear relationship to WAR. The result is a number that looks kind of like a fantasy dollar value--it's not, don’t go out and bid $65 for ARod in your fantasy league because he gets 65 points here, but the scale at least resembles fantasy dollar values.

If you read all of my stuff, you might be thinking “That’s awfully hypocritical, considering he doesn’t like unitless numbers like OPS or EQA”. Duly noted. However, the distinction as I see it is that WAR*7 has a direct, one-step relationship to a meaningful unit (just as Win Shares does, in theory). OPS needs an addition and a multiplication to be a decent run estimator, while EQA differences and ratios are unitless.

Third, I did not factor in defensive value, so everyone is assumed to be an average fielder at his position, thus making Hanley Ramirez the most valuable player in the NL, which I obviously don’t agree with.

Fourth, the system for classifying the producing organization of each player is not optimal. I credited each player to the franchise for which he made his major league debut. Obviously, players often change hands in the minors, and thus you might want to credit Grady Sizemore to the Nationals rather than the Indians. First major league organization is easy to do, though, and I don’t think it’s too much worse than going by signing organization. I believe that in Treder’s analysis, he tries to identify the organization most responsible for the player’s development, which could be either. This is a better approach, but I kept it simple here.

Fifth, the method of assigning players to teams was the same I used in my end of season stat reports, which means each player is credited to just one team even if played a significant amount with two teams (i.e. Saltalamacchia, whose name I find easier to spell than the star first baseman he was traded for).

Finally, I already slipped this in, but I only considered offensive players. So pitching, both on the team and produced by the system, has been completely ignored.

There are a number of different ways to look at the data, and what I am going to do is discuss a few of the interesting things, and then post a big chart at the bottom with all of the data.

Looking at the total (TOT) is not very interesting, because on the team level this just points out teams with talented position players. More interesting is the HG column, which measures homegrown talent retained by the team this year (actually, it includes anyone who made their major league debut with the team and played for So Sammy Sosa is considered homegrown for the Rangers despite the fact that he had been gone from the organization for more than fifteen years).

The leader in homegrown value was the Marlins, with 193 points. The next four teams on the list (Brewers, Phillies, Rockies, Braves) are all Neanderthal League outfits as well, while the Yankees top the AL, with a wide gap from those six teams back to the Indians.

The fact that the Yankees have a lot of homegrown value (James called it talent, and I may slip up and use that term too, but I want to stress that this is a measure of 2007 value and not talent) at first seems surprising, but consider Jorge Posada, Derek Jeter, and Robinson Cano all contributed significant value. Their total is boosted by the presence of Hideki Matsui, who in reality is a free agent signing, but here is treated as a Yankee product since he debuted in the majors with them.

There are a lot of good teams at the top of the homegrown list, but there are some pretty solid teams near the bottom too. One of those is the Cubs, with just 13 points of homegrown value (all contributed by Ryan Theriot). They are beat out by the Giants, though, whose seven points were all contributed by Pedro Feliz.

A logical jump is HG%, which is the percentage of total value produced by the team’s system. As you would expect, this has a strong correlation with the raw HG figure. Milwaukee led the way with 94% homegrown, with Florida, Minnesota, Colorado, and Philadelphia rounding out the top five. The Brewers got just 12 points from imported players (Johnny Estrada and Kevin Mench).

On the flip side, the Cubs (9%) and the Giants (7%) are the trailers. Teams as a whole got 55% of their offensive production from players they had developed.

Moving on, we have the “PROD” column, which measures the amount of value produced by the system. The leaders are the Indians with 278, just edging out the Marlins’ 273. I’ll look at the Tribe more closely in a bit, but Florida has produced 6 20 point players (~3 WAR), of which they retain(ed) five (Miguel Cabrera is now a goner). Only Edgar Renteria is gone. This may seem surprising considering the fire sales they have held, but a lot of the players they gave up in those trades were imports to begin with (Sheffield, Alou, Lowell), or no longer are around to produce any value.

The mean production is 161, pegging the Astros (162) and the Mets (158) as the most average organizations. The standard deviation is 66. I say this to set up that the z-scores range from -1.77 (Cubs) to +1.77 (Indians). With one exception, another half a standard deviation (-2.26) away from any other team. That team is the Giants.

When you see that the Cubs have only produced 44 points (around 6 WAR) of value, you can see that this is pretty bad. Their most notable contribution came from Brendan Harris (22), with the aforementioned Theriot next and just Corey Patterson and Ross Gload to chip in. But the Giants are on a whole different plane with a pitifutl 12. Only two San Fran products batted 300 teams with positive WAR this season--Feliz (7 points) and Yorvit Torrealba (5). I realize that this analysis overlooks a lot, especially pitchers, of which the Giants have a promising crop and a few good exiles out there. But it still strikes me as absurd that they rewarded Brian Sabean with a contract extension. Sabean got just two mediocre (for a playoff team) playoff teams out of four seasons of the greatest offensive force in baseball history, and his team has not been a real factor for a few years now. He has built an impossibly old team (although in his defense he has traded no prospects of offensive value to get it). You’re going to tell him “Nice job, we’d like another five years of this?”

Moving on, I have a column “%Retain”, which is the percentage of value produced by the system retained by the system (HG/PROD). The Rockies lead the way at 91%--only the Juans, Pierre and Uribe, are no longer members of the organization. They are followed by Philadelphia, Cincinnati, Detroit, and Milwaukee. The Tigers’ system has not produced much (61), but they retain 47 of it, and I doubt they’re too broken up about not having Juan Encarnacion, Frank Catalanotto, and Nook Logan. The major league average is 55%, which if you think about it makes sense--it has to be the same as the HG% on the league level.

On the other side of things, the White Sox stick out like a sore thumb with just 12% retention (the Padres are next at 26%). Magglio Ordonez, Carlos Lee, Aaron Rowand, Mike Cameron, and Frank Thomas are all 20 point players who have taken their services elsewhere by whatever means, while their most valuable retained product is a Japanese exception, Tadahito Iguchi. Josh Fields (10) is the most valuable true White Sock standing.

The “#” column gives the total number of players in the sample produced by each team, and “per #” is the per player average value produced by a system. The top three in producing players are Atlanta with 14, then Cleveland and Florida with 13. The average is eight and a third. The bottom three are San Francisco with 2, the Cubs with 3, and Detroit and Baltimore with five. Of course these lists are similar to the value produced lists.

In terms of value per farm product, Florida leads the way at 30, followed by the Yankees (27), Seattle, Philadelphia, and Colorado (26). Detroit has just 12 per player, Chicago 11, and San Francisco 6. So again, not only are the Cubs and Giants last in total value and players produced, the players that they have come up with are the least valuable.

You can play around with a lot of different combinations of the columns, but the last I will present is “Surplus”, which is the raw difference between Total and Production. A positive surplus means that the team had more value in 2007 than its system had produced. The average of course is zero, with the Astros (-4) being the closest. They, most notably, have lost Bobby Abreu, Luis Gonzalez, and Kenny Lofton, but they have also brought in Carlos Lee, Mark Loretta, and Mike Lamb with offsetting value (at least for 2007--the three that got away would have been a much bigger drain in, say, 2001).

The team with the biggest surplus (149) is Detroit, which has imported all of its notable offensive players except Curtis Granderson (Ordonez, Guillen, Polanco, Sheffield, Rodriguez). The flip side of the coin is their divisional foes, Cleveland. The Indians are short 124 points of value. You could make a pretty good team out of Indian exports (C: Josh Bard 1B: Sean Casey 2B: Brandon Phillips 3B: Kevin Kouzmanoff LF: Manny Ramirez CF: Coco Crisp RF: Brian Giles DH: Jim Thome). Even without a shortstop (and John McDonald is probably at least close to replacement level when you consider his defense), this team would have 157 points, which would rank it eighth in baseball (just ahead of the real Indians at 154).

Which feat do you find more impressive? That the Tigers have built a playoff contender on the basis of players brought in from elsewhere, or that the Indians have built a playoff team despite losing all of those players. It helps, I guess, that both teams have significant home grown pitching (on one hand Verlander, Zumaya, Bonderman, Robertson; on the other, Sabathia, Carmona, Betancourt).

Here’s a frivolous question for you: which team possessed the most value produced by another team? My off the cuff guess is the Yankees from the Mariners, on the strength of Alex Rodriguez. And that is indeed the answer. However, the second place finish is based on three players instead of just one--the Padres have 58 points of value produced by the Indians in Josh Bard, Kevin Kouzmanoff, and Brian Giles.

Here is the complete chart, which I sorted by total value produced:

Monday, December 10, 2007

Hitting by Position, 2007

This is a good once a year, mail-it-in type of post. As always, remember that we are dealing with just one year of data here, so it is not particularly significant, and any surprising findings should be viewed in that light. Nonetheless, it is a topic that I am always interested in and have fun looking at.

The data came from the Baseball Direct Scoreboard, which gets its data from STATS. I entered into the spreadsheet by hand, so I may have made a few errors, and those are my fault, not those of the Baseball Direct Scoreboard (I initially had Marlins shortstops down for 500+ walks, rather than the correct 53, and didn’t catch this until I looked at the position totals and shortstops came out as above average).

The first chart I have for you is the composite hitting by position. In addition to the standard positions, I like to look at 1B and DH together, and corner outfielders together. The “MLB” row are the MLB totals; they are not the same as the result you would get for all of the positions because of the way STATS compiles the data (pinch hitters don’t count, and there might be some other stuff). The “POS” row is the composite of the non-pitcher positions. “RC” is the basic version of ERP, as I did not include SB or CS data, and “PADJ” is the one-year offensive positional adjustment for each position, figured by dividing the position’s RG by the average position RG. “HPADJ” is the ten-year PADJ from 1992-2001 as a sort of baseline to compare against:

Again, I don’t want to read too much into the one-year data, but you can see that the range between positions is smaller than in the 1992-2001 period.

Insert obligatory comments on pitcher hitting and inflammatory comments about the Neanderthal League here. Pitchers created runs at a whopping 8% clip compared to position players. The top group of pitchers in terms of RAA compared to an average pitcher was the Cardinals, who took this coveted title for the second year in a row with a .195/.223/.238 line, +8 runs, three runs better than the Diamondbacks, Mets, and Dodgers. The worst was the Nationals at .112/.135/.138, -7 runs, just edging out the -6 turned in by the Astros, Giants, and Reds. Toronto pitchers had fun during interleague play, turning in 8 singles and a double in 22 PA for an above-total MLB average 5.1 RG.

I thought it would be fun to run a chart this year of the worst hitting teams at each position, which I have not done before. The best hitting teams at each position is boring, because the best players play almost all the time, and they usually play the same position. So it’s not at all interesting to report that the best hitting third base outfit was in the Bronx. However, the trailer list is a little more interesting, since bad players don’t usually take all 600 PAs themselves. So here are the worst at each position (no park or league adjustments here, BTW). The “RAA” column is against the average RG for the position in 2007:

Teams managed to overcome one bad position, as Cleveland and Arizona made the playoffs (although Arizona’s overall offense was poor; Cleveland was a bit above average). However, I can’t recommend being like the White Sox and having black holes at two positions.

A junk final thing I like to look at is the correlation, on a team level, between the long-term PADJ and the positional RGs. A positive correlation indicates that the team got their biggest offensive contributions from the left end of the defensive spectrum positions that you would expect; negative correlations indicate the opposite. Here are the team correlations (pitchers are not included for anyone and DH not for the NL). “AVG” is the average of the team figures, while “MLB” is the correlation between PADJ and RG for each position in the majors, individually:

I will show the data for three teams: the Astros, who had the strongest positive correlation; the Orioles, who had the weakest correlation; and the Yankees, who had the strongest negative correlation. The chart shows the position’s RG, the position’s ARG against the overall team RG (for the positions considered, i.e. no pitchers or NL DHs), and the 1992-2001 PADJ for each position as a benchmark:



The Astros repeat this honor; they led last year at +.91. You can see that they are weak up the middle, but get production out of the corners. Their team offensive spectrum goes 1B, LF, CF, RF, 3B, 2B, C, SS. Only CF is really misplaced relative to the defensive spectrum.



The Orioles’ correlation of +.03 is the lowest absolute correlation of any team, and you can see that by glancing at the numbers--they are over the map. The middle infielders are much better than the norm and their left fielders were the worst in baseball, but most of the other positions are fairly close to where one would expect.



The Yankees represent another repeat leader, as the correlation was the same -.34 in 2006. Six of their nine positions went the opposite way of what you would expect (i.e. you would expect first baseman to be above average; theirs were below average).

To end frivolously, I was impressed with the similarity between the production of the Cardinals’ center and right fielders. There may be better matches out there, but this one just happened to catch my eye. CF had 637 AB, RF 638. They each rapped out 170 hits, but RF had a little more power, winning in doubles (34-30) and homers (20-19). Triples went to center fielders 3-2, and the walk column was in their favor 56-49. Adding it all up, CF made 464 outs, RF 465. The center fielders created 87 runs and right fielders 86, giving them a 4.71 to 4.66 RG edge.

Raw data (stats by position for each team)

Monday, December 03, 2007

The Classes of 2003 and 2008

I felt that it was a very nice little coincidence this week when the two blockbuster trades both involved a young outfielder who first made a name for himself as a big high school prospect in the 2003 amateur draft. It is especially noteworthy to me, as I myself am a member of the class of 2003. Now of course I am not and never have been a major league prospect, but it is not hard to feel a bit of a connection to the players in the exact same age range as you, who are reaching adulthood at the same time as you. In basketball we had LeBron James to bear our standard, and in baseball, Delmon Young and Lastings Milledge were two of our very best prospects.

On the trades themselves, my opinion is not particularly interesting, since I have no special insight on the players and so many other voices have already weighed in. I think that the Rays-Twins trade was just a great baseball trade, one that will be fun to watch over the years to see who emerges on top. Were I running things in Tampa Bay, I would have found it very difficult to trade away a talent like Delmon Young, and I’m not usually one impressed by toolsy players with pitiful walk rates. If I had to guess, I’d say Matt Garza is able to provide more early value, but eventually the Twins win the deal. It’s a great trade, though, because eminently reasonable people can view it any number of ways.

The same will not be said of the Milledge deal. There is very little that can be said in defense of it, it seems; most comments that are not made by people with jaws on the floor tend to stress that Milledge is not that good of a prospect, and that the trade is not an all-time debacle. If that is the best that can be said for a deal, then it almost certainly shouldn’t have been made.

I have also seen the point raised that critics of the trade are overstating not only Milledge’s prospects but his trade value, and that Minaya obviously found that the value was only Schneider and Church. I have two problems with this argument, the first of which is simply that because Minaya felt that Schneider and Church was the most valuable package he could receive does not make it so.

The second is that the argument treats the relationship between a ballplayer and his general manager in the same manner as the relationship between a gallon of milk and a store manager. Even if we suppose that the Nationals’ package represents the extent of Milledge’s value, he’s not a commodity with an expiration date. He doesn’t need to be cashed in.

If I may make an even more ridiculous analogy, the relationship between the GM and the player is more like the relationship between me and my car. I have a car, and it has a certain resale value; which, in the case of my lovely early 1990s automobile, not a whole heckuva lot, but someone would take it at some price. If I sold the car, the best I might be able to get for it is $1,000. So if I do sell it for $1,000, have I made a good deal? After all, I got fair market value for it.

My answer is: depends. I was under no obligation to sell the car, so we have to consider other factors, like how much it will cost to replace my car. A better example is probably the stock market. If I sell a stock for $15, I have by definition received fair market value--there's not even a question about it. But if you were an investment advisor, would you just tell your clients to feel free to buy and sell stocks willy-nilly, just because by definition any stock transaction is a fair deal?

My point is that just because Milledge has X trade value doesn’t justify a decision to trade him for X. This is a trade for which it is very hard to see the upside for the Mets. If Brian Schneider and Ryan Church are going to outperform whoever the Mets’ other alternatives for their roster spots would have been, and doing so will make a significant impact on their fortunes, than I would submit that it is going to be a rough year in Queens.

Moving on, there is the Class of 2008. The potential Hall of Fame class, that is. As I have written before, I don’t really care about who goes into the HOF, because I don’t believe that the HOF has any capacity to honor the truly great players anymore (and “anymore” is not a new condition; the situation dates back to the 1970s at least) . I care a little bit, to about the same extent as I care about who wins NBA games. If Bert Blyleven is finally elected, it will still not be an honor to tell him that he is in the same class as Rube Marquard. As far as I am concerned, they can only dishonor him by waiting a dozen years before considering him worthy of standing aside Marquard.

If it was just Marquard, that would be on thing. But it’s not--it's Pop Haines, and Catfish Hunter, and Bob Lemon, and Chief Bender, and Dizzy Dean, and Jack Chesbro, and Lefty Gomez. None of whom should flatter Blyleven, or Tommy John for that matter, as company.

So I try to stay out of the HOF debates; while I like the “who was better than who” exercise as much as any baseball fan, it’s a lot more interesting to make your own lists, or follow along with something like the Hall of Merit or to just argue about players on a message board. So I’m not going to write an essay begging and pleading for the induction of Alan Trammell, Bert Blyleven, Tommy John, Goose Gossage, Mark McGwire, and Tim Raines, or bemoaning the fact that the BBWAA voters didn’t even give Lou Whitaker a second chance on the ballot. I’m going to write a sentence that does that, and move on with my life. Now excuse me; the Bobcats might be playing the Clippers right now.

Monday, November 19, 2007

Tangent Lines and Bill Kross

This is a math post with little baseball content and no baseball insight, so be forewarned.

In calculus, at least as far as I understand it, the tangent line is a line that intersects a point on a curve in the same direction as the curve, and the line has the same slope as exists on the curve at the point. That’s the best I can do--see this Wikipedia article for a better description.

Anyway, the tangent line is linear (it can be written as y = mx + b), and it shares the same slope as the line that it intersects. That means that near the point in question, it is just about the best linear approximation that you can get.

Where this ties into baseball is that if we have a non-linear function and want a linear approximation to it, the tangent line can be a shortcut that is easier and quicker than generating a line through some other technique (such as regression). Understanding how the tangent line works can also help us understand why non-linear baseball models have the linear approximations that they do.

First, let’s calculate a tangent line for a non-baseball problem. Suppose we have the line z = x^3, and we want a tangent line at the point x = 3. At x = 3, z = 3^3 = 27. The slope at x =3 can be found by first taking the derivative of z, which is z’ = 3x^2, so z’(3) = 3(3)^2 = 27.

We can write the line in the point-slope format as y - y1 = m(x - x1), where y1 and x1 are the base (x,y) point and m is the slope. So y - 27 = 27(x - 3). We can convert this to the common y = mx + b form to get y = 27x - 54.

At x =3, y = 27(3) - 54 = 27, which is exactly equal to z, as we know it should be. If we look at another x value close to 3, say 3.1, we get z = 29.791. We get y = 29.7. As you can see, they are pretty close. As we get further away, the linear approximation will perform worse, especially for functions with a steep slope.

Now, let’s talk about some of the baseball relationships where this is applicable. Clay Davenport used to publish a team version of EQR in which (RAW/LgRAW)^2 approximated the percentage to which the team R/PA exceeded the league average. There is also a linear version (which is the only one I have seen Clay publish in some time), in which the mapping is 2*(RAW/LgRAW) - 1.

Let’s call RAW/LgRAW “ARAW” for adjusted RAW. The two relationships we have are ARAW^2 and 2*ARAW - 1. Now suppose we work with the exponential function and find the tangent line at the league average point, where ARAW = 1 and the result of the formula = 1 (this is common sense, as a team with a RAW equal to the league average should score runs at a rate equal to the league average). The slope of ARAW^2 is 2*ARAW, which is 2*1 = 2 when ARAW = 1. So y - 1 = 2*(ARAW - 1), and y = 2*ARAW - 1. As you can see, that is the other Davenport
formula.

This is no surprise, as even if Davenport derived the relationship through a regression approach, we would expect the best fit to be about the same as the point at the league average, since most of the teams are tightly clustered around that point.

Another stat which follows the same relationship to runs is OPS. David Smyth (and perhaps others, but I recall seeing David write it) has pointed out that the square of relative OPS (not OPS+, but straight OPS/LgOPS) tracks runs, and Steve Mann wrote about the similar 2*(OPS/LgOPS) - 1 relationship eighteen years ago in The Baseball Superstats 1989.

The most interesting relationship, though, is the Pythagorean win estimator. I have written about this before on my website. Pyth can be written as:

WR = RR^z

Where WR is the win ratio (W/L), RR is the run ratio (R/RA), and z is the exponent (usually seen as z = 2). We know that for an average team, RR = WR = 1. The slope of the function is z*RR^(z - 1). If z =2, then it is just 2*RR, which is 2 when RR = 1. If z = 1.83 (another common value), than it would be 1.83*RR^.83, which is 1.83*RR when RR = 1.

We know that W% = WR/(WR + 1). We can therefore write this as a W% estimator as W% = (2*RR - 1)/(2*RR).

This method of estimating W% was discussed, informally, by Bill James in the 1984 Baseball Abstract. James said that if a team scored 10% more runs than their opponents, they should win 20% more games. He wrote that he had never tried it but it “should work”, and dubbed it “Double the Edge”. I have no idea whether Bill came up with this through similar mathematical logic to what you see here, or whether it was intuitive. With James, I’d believe either.

Anyway, the good thing about this estimator is that it caps W% at 1. However, it does not bottom out at zero--a RR of less than .5 results in a negative W%.

Ralph Caola, who has done a lot of work on run to win converters, emailed me after reading the article on my site and suggested that to solve this problem, one could use two equations: one when Run Ratio is greater than one, and one when Run Ratio is less than one. For the less than case, you could define W% as 1 - (2*OppRR - 1)/(2*OppRR), where OppRR is the opponents’ run ratio, RA/R. This way, reciprocal run ratios would produce complementary W%s, as we would intuitively expect (and as Pythagorean gives).

This way, reciprocal run ratios would produce complementary W%s, as we would intuitively expect (and as Pythagorean gives).

There are dozens of ways you can write those formulas, and Ralph settled on W% = (R-RA)/(R + RA + ABS(R-RA)) + .5.

And sure enough, the equation is more accurate and more theoretically sound if you use Caola’s insight. However, I have recently realized that Ralph was not the first one to uncover this formula. In fact, it has been in the public eye for over twenty years and little has been said about it. (I am not necessarily bemoaning this, because the only reason to use the linear approximations to Pythagorean is simplicity. They are not preferable. However, with the increased presence of sabermetric research all over the place, I am a bit surprised that Ralph and I seem to have been the only ones to play around with James’ Double the Edge).

In The Hidden Game of Baseball, there is a brief description of several run to win methods in Chapter 4. In a footnote, Palmer/Thorn write “About a year after Pete’s article [in SABR’s The National Pastime] appeared, Bill Kross, a Purdue professor, devised an elegant little formula that was not only simpler than the others, but also very nearly as accurate, erring only when run differentials were extreme (+/- 200 runs). If a team is outscored by its opponents, Kross predicts its winning percentage by dividing runs scored by two time runs allowed; if a team outscores its opponents, the formula becomes, 1 - RA/(2*R).”

Remember what I said about there being dozens of different ways to write the DTE formula? I am not going to go through the algebra here, but suffice it to say that the Kross formulas are one of the dozens. I don’t know if Mr. Kross developed those by linearizing the Pythagorean formula, or through some other technique, but there it is. These formulas are not a breakthrough in accuracy, be it empirical or theoretical, but they are quick and easy and do have a strong logical foundation, and can even be seen as offshoots of Pythagorean estimators.