Wednesday, October 04, 2006

Internet Baseball Awards: Manager of the Year

In this exciting series, I will share my votes for the Internet Baseball Awards, sponsored by Baseball Prospectus, and the thinking or lack thereof behind my choices.

Manager of the Year has always struck me as a silly award, because you can’t really evaluate a manager statistically. By convention you give it to the manager of the most surprising team, or to the manager of a team that is just really, really good. And of course which team you think are most surprising is biased based on your personal feelings before the season started, which may have been unfounded. Despite all of this, I will reluctantly choose in this manner as well.

In the American League, the big surprise team was the Tigers. The A’s managed to milk a .574 W% out of a .497 PW%, although it was the sub-.500 PW% that surprised me, not their division title. Some will hail the Twins as a surprise, but I picked them to play well. The East has a choke job in the Red Sox, the Blue Jays who did about as well as could be expected, and the DRays and Orioles are non-factors. Out West, the Angels, Rangers, and Mariners did about what I thought they would.

So I will go with the two “surprises” and the guy who managed the best team in the league:
AL MOY
1) Jim Leyland, DET
2) Ken Macha, OAK
3) Joe Torre, NYA

In the National League, the Marlins were pegged by me, foolishly in retrospect, to lose 100 games. There is a good chance that Joe Girardi will win manager of the year after being fired. The Astros and Reds surprised me, but Garner is not likely to garner a lot of support because they were pennant winners last year. The West played out like I thought it would, except I didn’t expect the Dodgers to end up in the postseason. So I have it:
NL MOY
1) Joe Girardi, FLA
2) Jerry Narron, CIN
3) Grady Little, LA

Tuesday, October 03, 2006

Quick Playoff Preview

This is not going to be extensive like those you will find elsewhere, but I just wanted to get on record with my picks (recognizing as always that a short series can easily turn any which way and not to take any prediction too seriously). And yes, my method of setting probabilities is flawed, as I treat PW% as the actual true W%, do not adjust for home/road, do not adjust for pitching matchups, etc. Anyway:
A’s v. Twins
A’s: 574, 529, 497 (W%, EW%, PW%), 4.81 R/G, 4.53 RA/G (park adjusted)
Twins: 593, 578, 553, 4.94, 4.22
Rooting for: A’s
As you can see, the Twins beat the A’s in every category, and the A’s component stats indicate a sub-.500 performance this year. While the A’s are my preference (like many sabermetricians, I have a soft spot for Billy Beane, but more importantly because they have my favorite player in baseball, Nick Swisher). One must remember that the Twins don’t have Liriano, who contributed to much of their success, but it’s still hard to pick against Minnesota.
P(Twins win game; Log5 based on PW%): 55.6%
P(Twins win series; above result, Binomial distribution): 60.4%

Dodgers v. Mets
Dodgers: 543, 546, 549, 5.27, 4.83
Mets: 599, 568, 572, 5.31, 4.65
Rooting for: Mets
Surprisingly, the Dodgers have below average defense, but their offense is strong enough to be equal, park-adjusted, with the vaunted Mets attack. This is a case where I am going to pick the upset: the Mets played poorly down the stretch, and really have a paper thin starting rotation (Glavine +7 RAA, El Duque -7, Trachsel -10, Maine +6).
P(Mets win game) = 52.3%
P(Mets win series) = 54.3%

Tigers v. Yankees
Tigers: 586, 597, 561, 5.23, 4.30
Yankees: 599, 608, 624, 5.86, 4.83
Rooting for: Yankees
The Yankees offense is by far the best in baseball, even more powerful then it was for most of the year with a full crew of Abreu, Sheffield, and Matsui. As Baseball Prospectus pointed out yesterday, the Tigers bench is putrid (Vance Wilson, Alexis Gomez, Ramon Santiago, Neifi Perez, Omar Infante). Their starters aren’t as good as everyone thinks they are; Verlander has a 4.63 eRA, Robertson 4.77, Rogers 4.35, Bonderman 4.51. This is a series the Yankees should win.
P(Yankees win game): 56.5%
P(Yankees win series): 62.1%

Cardinals v. Padres
Cardinals: 516, 513, 494, 4.90, 4.78
Padres: 543, 534, 555, 4.80, 4.46
Rooting for: Padres
This is a far cry from past Cardinal teams, limping down the stretch, with a surprisingly bad rotation (only Chris Carpenter ranks above average). The Padres have solid starters, a better bullpen, and almost as good of an offense. They are my pick.
P(Padres win game): 56.1%
P(Padres win series): 61.3%

I like the Dodgers over the Padres in the NLCS, the Yankees over the Twins, and the Yankees over the Dodgers in a World Series rematch of 1981…and 1978…and 1977…and….and 1941. But I am pulling for a Subway Series because I want to see the two best teams in baseball square off (although all of the NY hype for that is hard to stomach). But in the modern playoff system, that rarely happens.

Monday, October 02, 2006

2006 Park Factors

Herein I will present run and home run park factors for 2006 based on 5-year data (when applicable). I will provide a brief overview of the methodology--if you want a full description you can find it on my website or if you have any questions feel free to leave a comment. Basically, I first calculate a raw park factor (based on runs or home runs per game), accounting for the fact that the “true neutral” context includes a small percentage of games played in their home park as well as the road games (I do not, however, weight by the actual number of games played in each park). Then I take the average of that number and one, to account for the fact that only half of the games are actually played at homes (this means that the final result is meant to be applied to total statistics, not simply home statistics). Then I regress this figure towards one, with less weight given to one as the number of years of data we have for the park increases.

Here, are the PFs for 2006 (listed Run, HR):
ARI: 106, 106
ATL: 99, 98
BAL: 98, 103
BOS: 102, 94
CHA: 102, 113
CHN: 101, 107
CIN: 101, 108
CLE: 97, 93
COL: 112, 112
DET: 97, 94
FLA: 96, 93
HOU: 101, 105
KC: 100, 93
LA: 96, 104
LAA: 97, 94
MIL: 100, 103
MIN: 100, 95
NYA: 98, 101
NYN: 97, 95
OAK: 99, 100
PHI: 103, 108
PIT: 100, 95
SD: 94, 93
SEA: 96, 96
SF: 99, 90
STL: 99, 97
TB: 99, 98
TEX: 107, 108
TOR: 103, 107
WAS: 97, 94

There has been some talk about how Coors Field has played differently this year, often attributed to more aggressive use of the humidor. Coors Field did have its lowest raw home/road run ratio of the last five years--just 1.15, compared to a previous low of 1.24 in 2003. However, this still made it the second most offense-friendly park in the majors this year (Great American in Cincinnati was at 1.15 as well; Royals and Bank One were in that neighborhood as well. Side note: Yes, I know its not called Bank One anymore. Or Royals for that matter. I don’t care, I use the name I’m used to). Some people will argue that this is a fundamental change and that it should be handled differently then other parks, but I’m going to use the five-year average. If somebody could show me that it was a statistically significant difference, that would be one thing, but I haven’t seen evidence to that effect and anecdotes about the ball being moister or whatever the claim is don’t fly with me. However, I will give you the results by a couple alternative ways I could calculate the Coors PF (I certainly do not hold all of these as equally valid)
1) Treat 2006 as one year, don’t regress. Then the PFs would be 107, 108
2) Treat 2006 as one year, regress towards 1 like I would for another park. 104, 105
3) Treat 2006 as one year, but regress towards a historically normal Coors PF like 1.15 instead. 110, 111

If I was going to do something out of the ordinary, it would be option 3. As MGL has pointed out, parks shouldn’t be regressed towards 1. Each park should be regressed towards its own expected value, which we could determine based on a number of factors such as altitude, fence distance, surface type, foul territory, weather, outdoor/dome, etc. However, I have not undertaken the research that would be necessary to produce these custom values, and I don’t believe that anyone else has, at least not published and with a formula that can be used to combine all of the factors (obviously studies have been done into the more specific areas, but you would have to consider all of them in order to properly do what I am talking about).

So since we know that Coors Field has traditionally played as a very strong hitter’s park, and since we know that altitude is one of the largest factors in how a park should play (and that Coors is clearly at a very high altitude, i.e. very conducive to offense), it is silly to expect that the unseen mean we are regressing to for it is 1. That is what I have done in calculating all the parks, of course, but when we have five years of data, the differences are fairly negligible (If I didn’t regress the five year Coors results at all, I would get 114, 113 versus the 112, 112 I show above). For one year, though, it would definitely be an issue, especially for a park like Coors that we expect to be extreme, even if the balls are damp or whatever exactly it is they are supposed to be this year that they weren’t in the past.

Wednesday, September 27, 2006

Evaluating Pitcher W%, Pt. 2

As discussed in the first part of this series, an overlooked aspect of the traditional NW%/WAT approach is that it makes certain assumptions about how a team achieves its winning percentage (namely, all through the efforts of non-pitchers). So why not attempt to improve our methodology by using a more realistic model of W% causation?

I should note at this point that a lot of the ideas I am going to discuss were first published by Rob Wood in the August, 1999 edition of By the Numbers (see link to BTN archives on the right side of the page). While my results may not exactly match his, and my explanation is my own, it would be disingenuous to not acknowledge that he did this stuff first.

For any given team, our best assumption will be that their offense and defense are equally responsible for the team’s deviation from .500. Certainly this assumption will be wrong in some cases, worse then using the Oliver or Deane assumptions. But it will be correct more often and the overall error introduced by this assumption will be less then for others.

To keep things simple, I will assume that all defense is pitching. This is an obviously faulty assumption, but it will keep things workable, and again while this assumption will not always hold, it is better to assume that all pitching is defense then to assume that all deviation from .500 is the product of the offense. If one wanted to get even more precise then we are going to, they could introduce a correction for this.

Suppose we have a pitcher working for a team with Mate .540, who goes 15-10(.600 W%). His NW% and WAT under Oliver are .540 and +1. Under the Deane method, they are .565 and +1.63.

However, we are now going to assume that this team has pulled itself away from .500 through equal efforts by the offense and the pitching (excluding the pitcher in question of course, since we have removed his decisions from the rest of the team’s when calculating Mate). Using the Pythagorean theory, we can write this equation:
Mate = x^2/(x^2+(1/x)^2)

x is the percentage of league average runs the team must score (or inversely allow) in order to achieve a given W%. x can be solved for:
x = (Mate/(1-Mate))^.25
Thus, we expect a .540 team to score runs at 104.1% of the league average and allow runs at 96.1%. Once we know this, we can calculate the W% that we expect the team to have given only the non-average offense (since we are assuming that all defense is pitching and the other pitchers are irrelevant when evaluating our pitcher) to have a W% of x^2/(x^2+1), in this case .520.

We can also solve for the implied runs allowed ratio of our pitcher. He achieved a W% of .600 on a team with an offense that scored at 1.041% of average, so:
W% = x^2/(x^2 + y^2)
.6 = 1.041^2/(1.041^2 + y^2)

y can be solved for as:
y = x*sqrt((1-W%)/W%)
Or in this case, y = .85.

Now we know that by achieving a .600 W% for a team of this caliber, the pitcher’s performance was equivalent to allowing runs at 85% of the league average. To calculate his Neutral W%, we put him on a team with an average offense, and find that 1/(1+.85^2) = .581. That makes his WAT +2.03--significantly different then the Oliver and Deane estimates, because they (largely) assume that only the offense has caused the team to rise above .500, whereas we are assuming that it is a joint and balanced effort between the offense and the other members of the pitching staff.

We can generalize this for non-2 exponents as x = (Mate/(1-Mate))^(1/z), y as x*((1-W%)/W%)^(1/z), and NW% as 1/(1 + y^z), where z is the exponent we are using. But is any of this really necessary?

We found that a balanced .540 team would allow a .500 pitcher to be a .520 pitcher. If we simply calculate NW% for our pitcher as .600-.520+.500, as we did for Oliver, we find .580--pretty much equivalent to our convoluted Pythagorean approach. Not only that, but .520 is also equal to the average of .540 and .500. So can we just use this kind of approximation?

While I am a strong advocate of using methods that are theoretically sound across as many potential contexts as possible, practically we only care about the real range of major league teams, which I’ll just assume for the modern times are bounded between .250 and .750. If we find our expected W% for a .250 team with only the offense, it is .366. The average of .5 and .25 is .375, a difference of less then 3%. So it seems pretty safe to use this simplified assumption, and not screw around with all of the Pythagorean calculations. I should also note that there is a further error introduced when you eschew the calculation of the pitcher’s NW% through the Pythagorean approach as well. So there are two sources of error 1) is estimating the comparison level as the midway point between Mate and .500 and 2) is estimating that the pitcher’s NW% will be the same linear difference from .500 as his W% is from the estimate in part 1).

Again, I should note that this is the conclusion that Rob Wood came to, and also the same as Tango Tiger’s quick and dirty method linked below, in response to the first installment of this series. (As an aside, that is the problem with doing series in installments as I am wont to do, and at the same time having a few smart people read it. They figure out what you are doing, or what you should be doing, before you post it. I should either scrap the installment approach or not allow any readers). So, to summarize the quick approach:
NW% = W% - (Mate + .5)/2 + .5

How does this kind of approach change the standing of the historical pitchers discussed in the first installment? Well, Red Ruffing now has a NW% of .521 and +10.5 WAT. Still not in the league with many other Hall of Fame pitchers, but far from concluding that he was a true sub-.500 pitcher. Steve Carlton in 1972 now has a NW% of .846 and +12.8 WAT. Interestingly, he does better under this approach then the Deane approach, probably because the Deane approach looks at the percentage of possible improvement. But in Carlton’s case, the offense is, at least by our assumptions, so bad, that even an otherworldly performance can only do so much to raise the team’s fortunes.

In a third and final installment, which I promise will be posted by the end of the decade, I will look at replacement level and ask the question “Why even compare to Mate at all”, and perhaps throw in some other odds and ends.

Rob Wood in Aug 99 BTN(pdf)

Tango's blog entry

Thursday, August 24, 2006

Evaluating Pitcher Winning %, Pt. 1

How best to evaluate a pitcher’s W-L record? While it has plenty of contextual biases, one that it does not have is park/era, since W% always is .500 for the league as a whole. This makes pitcher win-loss record a fairly interesting thing to look at, at least on the career level.

But of course the biggest pollution is the quality of the team around him. So it only seems natural that for many years, would-be sabermetricians have compared a pitcher’s W% to that of his team. Usually, this comparison is done only after the pitcher in question’s decisions have been removed. The reasoning for this is that we do not want to compare the pitcher to a standard that he himself has contributed to. Anyway, Ted Oliver’s Weighted Rating System was the first such approach, and the one most commonly used:
Rating = (W% - Mate)*(W + L)

Where Mate, to borrow a designation from Rob Wood, is the W% of his teammates (TmW - W)/(TmW + TmL - W - L). Oliver’s rating gives a number of wins above what an average teammate would have achieved in the same number of decisions. We could also call this Wins Above Team as Total Baseball does.

A related question is what is the projected W% of this pitcher on an otherwise .500 team? I’ll call this Neutral W%, to use the same abbreviation but a different name then Bill Deane does, so that my general term won’t get confused with his specific one. For the Oliver approach:
NW% = W% - Mate + .500

If this is not intuitively obvious, consider a 20-10 pitcher on a .540 Mate team. His WAT is (.667 - .500)*(20 + 10) = +3.8. If he is 3.8 wins above average in 30 decisions, this implies that he is 3.8 wins better then 15-15, or 18.8-11.2. This is an equivalent W% of 18.8/30 = .627, the same result as .667-.540+.500.

What begins to become clear as you look at how the method works is that it assumes that a .500 pitcher on this team would have a .540 record. This means that all of the team’s deviation from .500 is attributed to the offense or fielders. This assumption is clearly wrong, at least for a randomly selected team--given a random team, we should assume that they are equally skilled on offense and defense. Obviously, in some cases this assumption will be dreadfully wrong--but it will be correct more often then assuming that EVERY team deviates from .500 only because of offense and one particular pitcher whom we isolate to calculate his WAT/NW%.

We can find some historical examples where the assumption of the Oliver method really causes problems. The most notorious case is that of Red Ruffing, who so far as I know is the only Hall of Fame starter with a W% worse then that of his teammates. For his career, Ruffing was 273-225(.548), while the rest of his team was .554. This is a .494 NW% and -3 WAT. As a side note, WAT is also equal to (NW%-.500)*(W+L).

Ruffing did pitch for Yankee teams with great offenses, but he also had mound teammates like Lefty Gomez, Johnny Allen, and Spud Chandler (at various times). In 1936, for example, Ruffing was 20-12(.625), while the rest of the team was .678, for -1.7 WAT and a .447 NW%. His team did score a whopping 1065 runs, but they also led the league with 731 runs allowed. An average pitcher in the 1936 AL (who would have a 5.67 RA), would figure to have only a .594 record if supported by New York’s 6.87 runs/game.

We’ll check in on Ruffing more as we go. Bill Deane, formerly a Senior Research Associate at the Hall of Fame, developed his own method to divorce a pitcher’s W% from that of his team. Deane’s insight was that the further above .500 a team’s W% was, the less margin there was to improve upon it. A .500 team could be bettered by .500; a .625 team only by .375. A bad team could be improved by even more. So Deane rated equally pitches who improved their teams by equal percentages of the potential margin.

A .550 pitcher on a .500 team improved his team by .050 out of a possible .500 (10%); so did a .460 pitcher on a .400 team (.060/.600 = 10%). Thus, they are each credited with the same .550 NW% (Deane used the term Normalized W% for this). If it is not clear why the normalized percentage should be .550 for each pitcher, it is because a .500 team has a .500 margin for improvement, and 10% of .500 is .050. Following this logic, Deane would up with these formulas for NW%:
If W% >= Mate:
NW% = (W% - Mate)/(2*(1 - Mate)) + .500
If W%< Mate:
NW% = .500 - (Mate - W%)/(2*Mate)

The second formula comes from the fact that on a .600 team, there is a .600 margin for lowering the W%; a .550 pitcher did this by 8.33%, so .0833*.5 = .042, for a NW% of .458. Total Baseball (unlike Thorn & Palmer’s earlier Hidden Game which used Oliver’s formula) used Deane’s methodology to calculate WAT. A poster child for considering the margin for improvement is Steve Carlton in 1972, who was 27-10(.730) for a team that was otherwise 32-87(.269). Under the Oliver methodology, this is a nearly impossible .961 NW% and +17.1 WAT. Using Deane’s approach, it is an .815 NW% and +11.7 WAT (still the highest since Lefty Grove in 1931).

How does Ruffing fair under this approach? Career-wise, since his W% was so close to Mate to begin with, not much changes--he now sports a .495 NW%(v. .494) and -2.7 WAT(v. -3). In 1936 he moves from .447 to .461 and from -1.7 to -1.3 WAT.

Thursday, August 10, 2006

Third and Third

As I write this, yesterday Oregon State, pretenders to the abbreviation of OSU, won the College World Series (yes, I really did write this in June). I figured it would be a good time to look back at the season of The OSU.

The Buckeyes finished third in the B10 regular season, crippled by a sweep in the heart of darkness. Northwestern shockingly was able to grab second after a horrific non-conference performance. Minnesota had a second consecutive year where they were not a major player in the race for the regular season title, but still qualified for the six-team tournament, which was filled out by Purdue and Illinois.

In the tournament, OSU beat Purdue and Northwestern but were tripped up in the winner’s bracket final by Minnesota and then lost the loser’s bracket final to those who shall not be named. They who shall not be named beat Minnesota two straight as the Gophers for the second straight year placed second in the tournament (to OSU in 2005). So those who shall not be named got the only B10 bid to the NCAA tournament.

While the Buckeyes fell short of a championship, they still had a solid season. Considering all games, OSU was second in W% at 37-21, .638 (those guys led at .672). But the Buckeyes paced the conference in EW%(.721; Minnesota was second at .627) and PW%(.728 with Minnesota second at .628). The Buckeyes also led in R/G(6.66; MSU second at 6.18) and RA/G(4.17; the bad guys second at 4.42). Northwestern’s W%, EW%, and PW% were .411(ninth of ten), .467(fifth), and .417(ninth), a simply bizarre combination for a second-place team. They were lucky that they did not have to face OSU, but even had they played the Bucks and been swept they would have finished in the first division.

With that, I will take a look at the individual performances of OSU players. Incidentally, all of the spreadsheets I used will be posted soon on my website if you are interested. Offensively, the Buckeyes were led by B10 MVP Ronnie Bourquin, the third baseman who was a second round pick to the Tigers. He narrowly missed the B10 triple crown, and hit 416/490/612 with 67 RC, a 12.2 RG(versus a conference average of 5.63, and +36 RAA. As you can see, his ISO was .196, but various scouting reports I saw before the draft said that he had power potential he had not shown in games. I have no trouble believing this, and can certainly understand why nobody on the collegiate level tried to mess with the form of a .400 hitter.

The Ohio offense was solid from top to bottom--sophomore centerfielder and leadoff man Matt Angle improved greatly, with a .449 OBA, 25-29 stealing, and +21. Sophomore catcher Eric Fryer was great again, with more power but less walks then Angle, resulting in nearly identical values(Angle created 54 runs and 9.2 per game; Fryer 54 and 9.2 per game). Senior captain and eighth round Oriole selection Jeddidiah Stephen finished his career at +16, 8.2, and his junior double play partner Jason Zoeller was second to him on the team in isolated power, +11 runs and 7.8 per game.

Junior Jacob Howell struggled through hamstring injuries, but hit a sizzling 402/448/500, 9.9, +15 RAA when able to play. The two weak spots in the lineup were Justin Miller, a freshman first baseman who started slowly but improved as the year went on, finishing at 4.3, -5. Junior rightfielder Wes Schirtzinger struggled greatly at the plate, with 257/321/296, 3.8, -11. The other hitters with over 100 PA were freshman OF/1B/P JB Shuck (6.1, +2) and DH Adam Schneider (4.7, -4).

The Buckeye pitching was solid again, tops in the B10 without a real standout. The ace was junior lefty Dan DeLucia with a 3.67 RA, +27 RAA, and 5.8 K per game, which may be why he went undrafted. Cory Luebke, a 22nd round pick of the Rangers as a draft eligible sophomore was 4.34 and +15. Freshman Jake Hale was the (relative) weak link at 4.92, +7. B10 Freshman of the Year JB Shuck probably looked better with traditional stats, as is 4.56 RA was a full two runs higher then his 2.51 ERA. Shuck, depending on your perspective, was victimized by his defense or had some mistakes obscured by the silly points of the earned run rule. His 4.52 eRA and .298 H/BIP lead me to the latter. But for a freshman, 79 innings and 12 runs above average is nothing to sneeze at.

There were really only four pitchers who got significant innings out of the pen. Rory Meister served as closer and had a 4.36 RA despite a 5.76 eRA. His control was very poor, walking 28 in 33 frames, but his H/BIP was a very high .382. Josh Barerra, a true freshman, had similar issues, walking 20 in 38 innings with a 5.68 RA and 7.16 eRA but a .429 hit rate. Both pitchers struck out a lot of batters and have shown evidence that they can be effective, but certainly need some polish. Trey Fausnaugh was pounded again with a 6.11 RA and 8.13 eRA. As were the other key relievers, he was victimized by a high hit per BIP rate at .409. Dan Barker was good again in 4 starts and 14 relief appearances, with a 4.15 RA and 3.57 eRA.

This was a fairly young team, but with only two (potentially three if Luebke was to sign with Texas) major losses, and a solid performance, it looks as if Ohio State will once again be a major player in the 2007 Big Ten race.

Thursday, July 13, 2006

Bafflement

I have always considered myself to be much more a fan of baseball in general then of any specific team. This is not the case for me in other professional sports. That is not to say I am not interested in NFL games that do not involve the Browns, but when the Browns were moved, my interest level in the NFL plummeted. I do not at all believe that this would be the case if the Indians were to dissipate.

I have always rooted for all Ohio teams, although since I have lived in the Cleveland and Columbus areas, I am more partial to those teams then to those from Cincinnati (of course, Columbus has just one pro franchise, and it's the only team from Ohio in the NHL, so there's really no conflict there at all. But if it comes down to Indians/Reds or Browns/Bengals, I definitely side with Cleveland).

Anyway, all of this personal rambling is to get the point that my dual status of 1) being a bigger fan of the sport in general then of any specific team and 2) the Reds' status as my second favorite, rather then favorite team is quite a fortunate thing today. Otherwise, I would be infuriated that the Reds traded away two everyday players, 26 years of age, and both rated as above average hitters for their position a year ago (Lopez by quite a bit, Kearns by the skin of his teeth), in exchange for a couple of relief pitchers, Royce Clayton, and Brendan Harris. It is simply unbelievable to me.

And what really makes it galling is that Jim Bowden, whose tenure to this point in Washington has been embarassingly bad, has completely stuck it to his old employers.

Tuesday, July 11, 2006

All-Stat Stat Check

I thought it would be a useful exercise, for myself at least, to run a quick check of the all-star break statistics. There are great, updated sabermetric stat resources updated at The Hardball Times and Baseball Prospectus, but I always like to figure my own stuff myself rather then look at somebody else’s work. But for the rest of you, using those sites may well be more edifying then reading what follows.

First, teams. I have Expected W%, based on Runs and Runs Allowed, and PW%, based on RC and RC Allowed. I have presented these in the order W%, EW%, PW%, and interspersed flippant comments. At various points I will use words like “luck” to describe variation from EW% or PW%. This usage is not meant to indicate that the differences are conclusively due to chance and not systematic factors:
DET(670, 662, 619)
Not surprisingly lucky, but (depending on your perspective) surprisingly still first in baseball in PW%
CHA(648, 612, 579)
Exceeding expectations again
BOS(616, 581, 591)
NYN(596, 579, 568)
Lead NL in all three W% types
NYA(581, 582, 607)
TOR(557, 544, 580)
STL(552, 517, 505)
Perhaps even more vulnerable then they appear to be in the standings.
MIN(547, 532, 499)
SD(545, 531, 548)
LA(523, 562, 554)
OAK(511, 482, 449)
Disappointing despite being in first place. I’m sure Mr. Beane is aware of this, in some form.
TEX(511, 524, 540)
COL(506, 515, 508)
CIN(506, 484, 506)
SF(506, 515, 505)
MIL(489, 417, 477)
LAA(489, 489, 526)
ARI(489, 476, 483)
SEA(483, 506, 483)
HOU(483, 468, 467)
CLE(460, 548, 567)
The Anti-White Sox. Again. I’ve seen it suggested that Wedge should be fired, with the underperforming Pythagorean expectation in a big way two years in a row cited as a reason. I don’t buy this for a second. The answer this year almost certainly lies in the run distribution, and unless somebody can give me hard evidence to the contrary, it’s darn silly to believe the manager has control over this.
PHI(460, 461, 443)
BAL(456, 431, 423)
ATL(449, 490, 475)
FLA(442, 487, 476)
TB(438, 412, 430)
WAS(422, 428, 442)
CHN(386, 387, 424)
Worst EW% in the NL
KC(356, 357, 357)
Equally bad by any measure, and last in baseball in all of them
PIT(333, 429, 421)
Last in PW% in the NL, in addition to real W%

Now onto the pitchers. I will give the top five and bottom five in each league in a few “payoff” categories--i.e. run-based stuff, not components. For reference, the AL as a whole is hitting .273/.336/.437 with 5.05 runs per game; the NL is at .265/.330/.425 with 4.77. I used 75 innings as the qualification standard:
AL RA LEADERS:
Liriano, MIN (1.84)
Halladay, TOR (3.07)
Santana, MIN (3.16)
Verlander, DET (3.19)
Contreras, CHA (3.46)
AL eRA LEADERS:
Liriano, MIN (2.43)
Lackey, LAA (2.76)
Santana, MIN (3.12)
Halladay, TOR (3.24)
Bonderman, DET (3.29)
AL RAR LEADERS:
Halladay, TOR (+47)
Santana, MIN (+46)
Liriano, MIN (+44)
Zito, OAK (+39)
Verlander, DET (+38)
I guess Santana would get my mid-season Cy Young vote, but Liriano and Halladay are having superb seasons as well.
AL RA TRAILERS:
Silva, MIN (7.70)
McClung, TB (7.41)
Lopez, BAL (7.20)
Weaver, LAA (6.94)
Johnson, CLE/BOS (6.88)
AL eRA TRAILERS:
Silva, MIN (7.22)
Johnson, CLE/BOS (6.75)
McClung, TB (6.72)
Lopez, BAL (6.57)
Radke, MIN (6.56)
The greatness of having the Santana/Liriano combo is lessened by having two of the league’s worst performers.
AL RAR TRAILERS:
Silva, MIN (-14)
Lopez, BAL (-11)
McClung, TB (-10)
Weaver, LAA (-6)
Johnson, CLE/BOS (-5)
Now some lists that will be presented with leaders and trailers back-to-back, since they’re not really “good/bad” indicators:
HIGHEST RA/eRA ratio, AL:
Lackey, LAA (3.49/2.76)
Santana, LAA (4.52/3.62)
Johnson, NYA (5.68/4.61)
Vazquez, CHA (5.33/4.64)
Wakefield, BOS (4.61/4.11)
LOWEST RA/eRA ratio, AL:
Liriano, MIN (1.84/2.43)
Lilly, TOR (4.54/5.32)
Verlander, DET (3.19/3.73)
Kazmir, TB (3.67/4.30)
Radke, MIN (5.65/6.56)
HIGHEST $H, AL:
Johnson, CLE/BOS (.354)
Millwood, TEX (.345)
Radke, MIN (.344)
Silva, MIN (.342)
Weaver, LAA (.340)
LOWEST $H, AL:
Lackey, LAA (.237)
Beckett, BOS (.253)
Elarton, KC (.259)
Zito, OAK (.261)
Verlander, DET (.262)
NL RA LEADERS:
Webb, ARI (2.91)
Penny, LA (2.91)
Schmidt, SF (3.00)
Johnson, FLA (3.06)
Young, SD (3.30)
NL eRA LEADERS:
Schmidt, SF (3.17)
Martinez, NYN (3.24)
Webb, ARI (3.49)
Young, SD (3.61)
Johnson, FLA (3.62)
What a nice surprise Josh Johnson has been. I think Texas would like to have Chris Young back right about now too (although there are no park factors considered here, and Petco does cut into runs by about 5%).
NL RAR LEADERS:
Webb, ARI (+47)
Schmidt, SF (+42)
Arroyo, CIN (+37)
Penny, LA (+37)
Capuano, MIL (+36)
I guess that makes Webb or Schmidt the Cy Young choice.
NL RA TRAILERS:
Perez, PIT (7.58)
Moehler, FLA (7.34)
Madson, PHI (6.66)
Claussen, CIN (6.55)
Suppan, STL (6.52)
NL eRA TRAILERS:
Moehler, FLA (6.99)
Perez, PIT (6.87)
Madson, PHI (6.85)
Sosa, ATL (6.64)
Herandez, WAS (6.60)
What the heck happened to Oliver Perez? His strikeout rate has dropped now as well to 7.2
NL RAR TRAILERS:
Perez, PIT (-14)
Moehler, FLA (-12)
Hernandez, WAS (-7)
Madson, PHI (-7)
Suppan, STL (-6)
HIGHEST RA/eRA ratio, NL:
Bucholz, HOU (5.43/4.20)
Cain, SF (5.53/4.47)
Martinez, NYN (3.91/3.24)
Olsen, FLA (4.84/4.27)
O’Connor, WAS (4.80/4.24)
LOWEST RA/eRA ratio, NL:
Penny, LA (2.91/3.76)
Oswalt, HOU (3.38/4.19)
Glavine, NYN (3.86/4.70)
Maholm, PIT (5.29/6.41)
Webb, ARI (2.91/3.49)
HIGHEST $H, NL:
Maholm, PIT (.351)
Moehler, FLA (.348)
Pettitte, HOU (.346)
Madson, PHI (.341)
Snell, PIT (.336)

Let’s look at hitters now (200 PA needed to qualify):
AL BA LEADERS:
Mauer, MIN (.378)
Johnson, TOR (.365)
Jeter, NYA (.345)
Suzuki, SEA (.343)
DeRosa, TEX (.332)
AL OBA LEADERS:
Hafner, CLE (.457)
Mauer, MIN (.451)
Ramirez, BOS (.436)
Catalanatto, TOR (.431)
Johnson, TOR (.424)
AL SLG LEADERS:
Thome, CHA (.651)
Hafner, CLE (.650)
Dye, CHA (.646)
Thames, DET (.634)
Ramirez, BOS (.615)
AL RC LEADERS:
Hafner, CLE (79)
Ortiz, BOS (75)
Ramirez, BOS (74)
Thome, CHA (73)
Wells, TOR (69)
AL RG LEADERS:
Hafner, CLE (10.24)
Ramirez, BOS (9.15)
Thome, CHA (9.04)
Mauer, MIN (8.86)
Dye, CHA (8.67)
AL SEC LEADERS:
Giambi, NYA (.585)
Hafner, CLE (.577)
Thome, CHA (.547)
Ramirez, BOS (.540)
Ortiz, BOS (.508)
AL BA TRAILERS:
Anderson, CHA (.192)
Lee, TB (.197)
Reed, SEA (.217)
Sexson, SEA (.218)
Ellis, OAK (.219)
AL OBA TRAILERS:
Reed, SEA (.256)
Hall, TB (.258)
Uribe, CHA (.259)
Berroa, KC (.265)
Ellis, OAK (.271)
I really don’t understand the Hall and Hendrickson for Seo and Navarro trade at all. Sure, Hendrickson is an upgrade over Seo, at least for the present, but you give up a pretty good young catcher and get an older catcher whose never done much. But it makes sense to Brian Sabean and his acolytes, apparently
AL SLG TRAILERS:
Lee, TB (.291)
Ellis, OAK (.311)
Kendall, OAK (.314)
Anderson, CHA (.324)
Ford, MIN (.324)
A first baseman last in the league in slugging is never a good thing. At least Jason Kendall has a homer this year.
AL RG TRAILERS:
Lee, TB (2.60)
Ellis, OAK (2.79)
Anderson, CHA (2.91)
Berroa, KC (2.93)
Reed, SEA (3.01)
AL SEC TRAILERS:
Polanco, DET (.110)
Berroa, KC (.129)
Kendall, OAK (.138)
Loretta, BOS (.142)
Betancourt, SEA (.144)
Now the Neanderthal League:
NL BA LEADERS:
Garciaparra, LA (.358)
Sanchez, PIT (.358)
McCann, ATL (.343)
Holliday, COL (.337)
Lamb, HOU (.337)
NL OBA LEADERS:
Bonds, SF (.460)
Abreu, PHI (.451)
Pujols, STL (.432)
Cabrera, FLA (.428)
Helton, COL (.421)
Neither league’s OBA leader is in the All-Star game.
NL SLG LEADERS:
Pujols, STL (.703)
Berkman, HOU (.607)
Beltran, NYN (.606)
Holliday, COL (.587)
Howard, PHI (.582)
NL RC LEADERS:
Wright, NYN (72)
Pujols, STL (71)
Cabrera, FLA (70)
Berkman, HOU (69)
Bay, PIT (68)
NL RG LEADERS:
Pujols, STL (10.27)
Bonds, SF (8.55)
Garciaparra, LA (8.47)
Berkman, HOU (8.46)
Cabrera, FLA (8.30)
NL SEC LEADERS:
Bonds, SF (.640)
Pujols, STL (.590)
Dunn, CIN (.519)
Ensberg, HOU (.511)
Beltran, NYN (.509)
NL BA TRAILERS:
Lane, HOU (.205)
Barmes, COL (.208)
Guillen, WAS (.211)
Abercrombie, FLA (.221)
Molina, STL (.222)
NL OBA TRAILERS:
Barmes, COL (.234)
Guillen, WAS (.254)
Molina, STL (.256)
Castilla, SD (.259)
Burnitz, PIT (.269)
NL SLG TRAILERS:
Ausmus, HOU (.299)
Taveras, HOU (.308)
Schneider, WAS (.311)
Barmes, COL (.318)
Everett, HOU (.319)
Houston cannot expect to contend again with their sorry offense. Of course Berkman, Lamb, and Ensberg are having pretty good years, but when you have three black holes playing every day, you let that go to waste.
NL RG TRAILERS:
Barmes, COL (2.23)
Molina, STL (2.43)
Castilla, SD (2.57)
Everett, HOU (2.84)
Cedeno, CHN (2.90)
NL SEC TRAILERS:
Cedeno, CHN (.117)
Taveras, HOU (.118)
Eckstein, STL (.122)
Castilla, SD (.126)
Pierre, CHN (.135)
I didn’t include steals in Secondary Average here, so Pierre and Taveras might not be in the bottom five if you include them.

Now let’s look at the top and bottom three players, ranked by PRAR, at each position, for both leagues combined (although I will list the best and worst in each league if they are not among the extreme three).
CATCHERS:
Mauer, MIN (+39)
Martinez, CLE (+27)
Barrett, CHN (+23)
Hall, TB (0)
Ausmus, HOU (-1)
Molina, STL (-5)
Incredibly, A.J. Pierzynski, who is largely reviled, and is sixth among AL catchers in this category, is in the All-Star Game. Martinez’ throwing has been horrible this year, but he is still a great hitter for the position.
FIRST BASE/DH:
Hafner, CLE (+45)
Pujols, STL (+42)
Thome, CHA (+37)
Sexson, SEA (-4)
Everett, SEA (-6)
Lee, TB (-12)
Hafner leads the world in PRAR, and yet he doesn’t get that much respect on a national level. Pronk coming into this year had a career line of 293/388/556 while David Ortiz was 282/366/534. I realize Ortiz is the “clutch God”, is charismatic, and was a member of a world championship team, so I’m not surprised that he is more famous, but I think it is disproportionately so. And I’m not trying to put down Big Papi--he was next on this list at +33. I’d love to have either of them in my lineup. Adrian Gonzalez is last in the NL at +6.
SECOND BASE:
Utley, PHI (+33)
Uggla, FLA (+29)
Hall, MIL (+25)
Kennedy, LAA (+2)
Polanco, DET (+1)
Ellis, OAK (-4)
Perhaps the greatest advantage for the Neanderthals is at second base. You have to go all the way down to seventh on the list to find the AL leader, Brain Roberts, at +15. Aaron Miles at +3 is the worst in the NL.
THIRD BASE:
Cabrera, FLA (+40)
Wright, NYN (+38)
Rolen, STL (+31)
Bell, PHI (+5)
Beltre, SEA (+4)
Boone, CLE (-2)
ARod for all of the grief he gets, leads the AL at +28 (Chipper is ahead of him in addition to the three above). Aaron Boone continues to stink, and while Andy Marte has not had the best season at AAA, it’s time to let the guy play in the majors. It can’t get much worse then it already is.
SHORTSTOP:
Jeter, NYA (+35)
Reyes, NYN (+34)
Tejada, BAL (+31)
Everett, HOU (-1)
Berroa, KC (-2)
Barmes, COL (-7)
Remember the SI article in the 1996 baseball preview about how New York would have the best pair of shortstops in either league with Jeter and Rey Ordonez? They were ten years ahead of their time. Reyes, while his style is still not the kind of baseball I prefer, has been a favorite of mine this year as I used my second round fantasy pick on him, which prompted a comment of “What?” from another league member. I still wouldn’t pick him in the second round for a real team, but he still has time to change that, and he’s a pretty good player at this moment.
CORNER OUTFIELD:
Ramirez, BOS (+40)
Berkman, HOU (+37)
Dye, CHA (+34)
Abreu, PHI (+34)
Holliday, COL (+31)
Bay, PIT (+31)
Francouer, ATL (0)
Markakis, BAL (0)
Burnitz, PIT (-1)
Guillen, WAS (-4)
Ford, MIN (-5)
Abercrombie, FLA (-6)
I didn’t bother splitting them up into left/right, so I listed the top and bottom six. Jeff Francouer has eight walks in 369 at bats. JC Bradbury at Sabernomics has started a tracker to follow Francouer’s quest to make the most outs in a single season. He is simply not a good player, at least not at this point in his development.
CENTER FIELD:
Beltran, NYN (+37)
Wells, TOR (+35)
Matthews, TEX (+28)
Kotsay, OAK (-2)
Taveras, HOU (-4)
Anderson, CHA (-5)
Weren’t the New York fans booing Beltran at the beginning of the year too? Good grief, you’re the best in the league at your position and you get booed. If I had MVP votes, they would go to Hafner and Pujols, although Joe Mauer could make a good case depending on how his defense is (I’m not a big defensive stats maven and I’m not looking it up right now).