For some reason this winter I was struck with nostalgia for 1994 Topps, and decided that it would be a good off-season project to collect the complete set. I’d never collected a complete current set when I was interested in cards – I didn’t have the resources, and I preferred to get a variety of different sets (although as I said, 1994 Topps was my favorite based on how represented it was among my cards) or to buy Ken Griffey cards. Yes, my year of baseball card obsession corresponded with thinking that Ken Griffey was the best and coolest player. Not that I have anything against Griffey, but in retrospect, it now seems like a lot of wasted time and money that could have been spent on Barry Bonds or Rickey Henderson cards. Within a year I was getting seriously into statistics and discovering that Bonds, not Griffey, was the best player in baseball, but by then I wasn’t buying a lot of cards.
I say I never collected a complete “current” set because I do in fact have a complete set of 1991 Donruss, or extremely close to it. They are ugly cards, to be sure, but we used to buy boxes of them dirt cheap at Big Lots. I also have to be pretty close on 1990 and 1992 Topps, although I never sat down and tried to inventory what I had.
Of course it would have been easier and cheaper to just buy the complete set, especially for Topps, but I decided it would be more fun to start with my old collection, buy some unopened packs, and buy singles to fill in whatever gaps were left. So that’s what I did, and it was more fun and more expensive. I felt conflicted about opening previously sealed boxes of twenty-six year old cards; on the one hand, it felt like squandering the last of a non-renewable resource. On the other, these poor cards had been stuck inside for two and a half decades. Time for them to get out and be admired like they were made for.
There was a complication, though. I believe 1994 Topps was the first Topps set to be glossed on both sides, and they get stuck together. I was usually able to get them apart without too much damage, but there were a handful of unfortunate incidents, although I think Mike Mussina was the biggest name for which I ended up with a severely defaced card. At least by 1994 there was no disgusting gum enclosed.
There were frustrations -- I ended up with a lot of extras, naturally, but not as many of some of the cards that I now consider most desirable (Barry Bonds, Rickey Henderson, Roger Clemens, Jeff Bagwell) as I would have liked. As always seems to be the case, Alan Trammell and Lou Whitaker stuck together – I got zero of the former and only one of the latter. There were amusing coincidences as well. In one pack I got a Lee Smith and a Trevor Hoffman back-to-back – the man who contemporaneously held the career saves record, and the man who would break it. When I was opening a pack during commercials while watching the Super Bowl, I pulled a Pat Mahomes card.
One thing that struck me as I was putting the set together was that these were “my” cards, but not “my” players. 1994 Topps is and always will be the consummate ideal of a baseball card in my mind, but of course they represent the players of the 1993 season, in which I cared zilch about baseball. The 1995 or 1996 set would much better capture the universe of baseball players as I came to love it, but the cards themselves would not hold the same nostalgia value.
The other thing that struck me after I had mostly completed the set in the spring was how the baseball experience of 1994 would reverberate to that of 2020, the only two seasons of my time as a fan that have been catastrophically shortened (144 games in 1995 sounds real good right about now).
With that, I will write about the cards themselves. On the back, the second picture takes up about a quarter of the card, but leaves plenty of room for a statistical record that is probably better than what was published in contemporary Who’s Who In Baseball. You get the demographics (height/weight, bats/throws, draft status, how acquired by current team, DOB/birthplace, and current home), plus decent stats for the entirety of the player’s major league career (and if you don’t believe that, look at how they had to squeeze it in on Nolan Ryan’s final card). For pitchers, the statistical categories are G, IP, W, L, R, ER, SO, BB, GS, CG, SHO, SV, and ERA – weird order, but total runs allowed? Can’t argue with that. For hitters: G, AB, R, H, 2B, 3B, HR, RBI, SB, SLG, BB, SO, AVG. For 1994, that’s pretty good.
The set is not marred by inserts – there are a number of “special” cards, but none that are inserts – they all are standard numbered, not irritants designed to make it impossible to complete a set. Unfortunately, these special cards are the weakness of the set.
1. 1993 Draft Picks
At first, this should make you excited, because the #1 overall pick in 1993 was the greatest player ever to be taken #1 overall, and one of the greatest to be drafted period, Alex Rodriguez. But he is nowhere to be seen, nor are most of the other best players drafted in the first round (Brian Anderson, Chris Carpenter, Darren Dreifort, Torii Hunter, Derrek Lee, Trot Nixon, Jason Varitek), with Billy Wagner being the exception. The highest draft pick represented is Wayne Gomes (#4). I assume that since the draft picks weren’t MLBPA members, they were able to cut their own deals, which explains why they’re absent, but it’s still a disappointing lot, and there are about 35 cards wasted on them.
2. Prospects
These cards feature four players to a card, which means you get only a tiny picture of their faces. It’s typically one player from AAA, AA, A, and a draftee. There’s one card for each position, and it’s worth listing them out:
C (#686): Chris Howard, Carlos Delgado, Jason Kendall, Paul Bako
1B (#448): Greg Pirkl, Roberto Petagine, DJ Boston, Shawn Wooten
2B (#527): Norberto Martin, Ruben Santana, Jason Hardtke, Chris Sexton
3B (#369): Luis Ortiz, David Bell, Jason Giambi, George Arias
SS (#158): Orlando Miller, Brandon Wilson, Derek Jeter, Mike Neal
OF (#79): Billy Masse, Stanton Cameron, Tim Clark, Craig McClure
OF (#237): Curtis Pride, Shawn Green, Mark Sweeney, Eddie Davis
OF (#616): Eddie Zambrano, Glen Murray, Chad Mottola, Jermaine Allensworth
SP (#316): Chad Ogea, Duff Brumley, Terrell Wade, Chris Michalak
RP (#713): Todd Williams, Ron Watson, Kirk Bullinger, Mike Welch
Similar to the draft picks, I’m not sure how many players the powers that were at Topps at actually had the ability to choose for these cards. That being said, if this is the success rate, why bother? This did produce the first Topps card for Derek Jeter (I’m not going to wade through the excruciating minutia to try to figure out if it qualifies as his Rookie card or not), but sharing a card with the likes of Brandon Wilson and Mike Neal is not the stuff 1952 Topps Mickey Mantles are made out of. I’d say they hit a homer run with catcher, and the rest of these are just sad. How could you pick twelve outfielders and only have Shawn Green to show for it? And I for one am shocked that the relief pitcher prospects didn’t pan out.
3. Coming Attractions
These cards, the last 28 regular cards in the set (minus the Series 2 checklists that close it out; 4 of the 792 cards are devoted to checklists), feature two prospects from each team. These guys were all supposed to be close to the majors, and so there are more familiar names; I’m not going to list them all, but I’ll give you the top names to give you a flavor. Excluded from this list are the two Braves, arguably the two best players in the group on the same card, a definite success – Chipper Jones (the best player represented by leaps and bounds) and Ryan Klesko.
Carl Everett, Bob Hamelin, Scott Hatteberg, Raul Mondesi, Troy O’Neal, Bill Pulsipher, Steve Trachsel, Rondell White
4. Future Stars
These cards are sprinkled in, one for each team, showing a player pretty close to the majors. There are a couple big names, and as a kid this was my favorite card in the set (along with Ken Griffey, of course). I had multiple copies, and I also happened to pull a bunch of them when opening old packs:

Unfortunately, I don’t consider the design of these to be so cool anymore, with the “futuristic” border. The pixelization of everything in the background of the shot other than the player is not a horrible idea, but as an adult I’d just prefer a nice picture. The future stars were:
Paul Carey (BAL)
Frank Rodriguez (BOS)
Justin Thompson (DET)
Domingo Jean (NYA)
Alex Gonzalez (TOR)
Scott Ruffcorn (CHA)
Manny Ramirez (CLE)
Billy Brewer (KC)
Jose Valentin (MIL)
Rich Becker (MIN)
Garrett Anderson (CAL)
Eric Helfand (OAK)
Tim Davis (SEA)
Benji Gil (TEX)
Javy Lopez (ATL)
Nigel Wilson (FLA)
Cliff Floyd (MON)
Butch Huskey (NYN)
Tony Longmire (PHI)
Matt Walbeck (CHN)
Calvin Reese (CIN) (yes, the back of the card does mention “Pokey”)
Todd Jones (HOU)
Danny Miceli (PIT)
Tripp Cromer (STL)
Mark Thompson (COL)
Billy Ashley (LA)
Ray McDavid (SD)
Salomon Torres (SF)
The weirdest thing for me about grouping 1994 teams by division is the Tigers in the AL East.
5. 1993 Topps All-Stars
Apparently these are Topps’ own picks for 1994 All-Stars in each league. They appear in a run to close out the first series (at least until the checklists gum things up). The design of the cards is unremarkable, but what’s more interesting are the stats on the back – they split the players’ stats by pre- and post-All Star Break, even though they are full season selections, and the stats are different – they show Total Bases and OBP, but not SLG which are on the regular cards. Who knew it was easier to pick all-stars than prospects?
C: Mike Piazza, Mike Stanley
1B: Fred McGriff, Frank Thomas
2B: Robbie Thompson, Roberto Alomar
3B: Matt Williams, Wade Boggs
SS: Jeff Blauser, Cal Ripken
LF: Barry Bonds, Albert Belle
CF: Lenny Dykstra, Ken Griffey
RF: David Justice, Juan Gonzalez
SP: Greg Maddux, Jack McDowell
RP: Randy Myers, Jeff Montgomery
6. Measures of Greatness
These feature one active player of historical stature and compare their statistics to those of the average Hall of Famer and one particular Hall of Famer at their position (this statistical comparison is shown in Bill James’ seasonal notation, per about 158 games, which I assume was chosen since it's the average of 154 and 162). The pairings can be entertaining, though:
C: Darren Daulton/Roy Campanella
In fairness, this wasn’t a great time for great catchers – Carlton Fisk and Gary Carter had retired, Mike Piazza and Ivan Rodriguez were too young for this kind of company. This did not age well.
1B: Frank Thomas/Jimmie Foxx
2B: Ryne Sandberg/Rogers Hornsby
3B: Wade Boggs/Brooks Robinson
I’d make fun of this one, but there was still an extreme dearth of HOF third baseman at this point; Mike Schmidt wasn’t in the Hall yet and George Brett had just retired. Eddie Mathews is not a good comp for Wade Boggs at the plate stylistically, so the pickings are slim. There still hasn’t been a third baseman who really comps to Boggs.
SS: Cal Ripken/Luis Aparicio
Same problem, although Ernie Banks would be a much better fit than Aparicio. Here's the back of the card as an example of these:
OF: Barry Bonds/Willie Mays
This comparison looks even better now than it did then.
OF: Ken Griffey/Stan Musial
This is a bad comp (Mays would be the lazy one), or at least would be a bad comp were it not for the fact that both hail from Donora, Pennsylvania. Does the card point this out? Nope.
OF: Kirby Puckett/Joe DiMaggio
Really? Not Duke Snider or something?
DH: Paul Molitor/Roberto Clemente
I’ll cut them some slack since Molitor was multi-positional and closing in on 3,000 hits. Oddly, pitchers were omitted from Measures of Greatness.
Wednesday, August 19, 2020
1994 Topps, pt. 2
Wednesday, August 05, 2020
1994 Topps, pt. 1
If you asked me today what about baseball interested me most (besides the basics like which team is going to win the World Series or how the Indians are going to do this season), I would say “sabermetrics and scorekeeping”. That answer has probably been the same since 1996 or so. If you’d asked me in 1995 I would have said “statistics” rather than sabermetrics, and scorekeeping wasn’t on the radar. But if you’d asked me in 1994, statistics would have been second, but a fairly distant second. The answer would have been “baseball cards”.
I was always interested in facts and figures, so it’s no surprise that when I became an overnight baseball fanatic, I was caught up in lists of pennant winners and ERA leaders and the like, and this led me down a path to sabermetrics. I think my early fascination with baseball cards comes from already having collected football and basketball cards (which in turn came from an innate desire to collect things), so it was natural for interest in a sport to be followed closely by an interest in cards. Of course, just as is the case for statistics, baseball is the sport for card collecting, or at least certainly was circa 1994.
I am not a historian of baseball cards, so take the discussion that follows with a grain of salt. It’s written by the seat of my pants, and I could have done so much more accurately when I was nine. But 1994 seems to represent the zenith of the roller coaster history of card manufacturers that took the hobby from a Topps monopoly to five major players and eventually right back to a Topps monopoly. In 1994, the five majors were cranking out multiple sets, aimed at different levels of consumer. In retrospect it seems like an obvious bubble, in the way that bubbles usually do but only after the fact. Who was the audience truly demanding a super-premium card set?
Yet most of the majors had a base set (Topps, Fleer, Donruss, Upper Deck, Score); a premium set (Topps Stadium Club, Fleer Ultra, Leaf, Upper Deck SP, Pinnacle); a super-premium set (Topps Finest, Fleer Flair, Leaf Limited); plus the other sets which included Donruss’ Triple Play and Studio, Topps’ Bowman, and Upper Deck Collector’s Choice. The strike would help deal a blow to the insanity, but it seems like a market that was already ripe for a correction.
I try to actively avoid thinking that baseball or anything else peaked when I first fell in love with it, as so many people do. But I will always maintain that 1994 had the best cards (not that I know anything about post-1995 cards), the perfect combination of an increase in production values (gloss on both sides, although I will have more to say on that later) but when you could still collect a base set loaded with commons without a million insert cards. And aesthetically? They were (mostly) beauties.
Super-premium cards were inherently ridiculous, but I don’t think any looked better than 1994 Fleer Flair with its regal names and thick stock:
On the premium side, all of the offerings from the majors were memorable. Topps Stadium Club was the worst, but there’s something delightfully 90s about the names on the front of the card, with the lower case first name and the all caps surname straight out of the labelmaker your dad had stashed away in the basement. Although we’re not going to talk about what the backs of these cards looked like:
Upper Deck always put out a classy card (although I will admit that these rank well behind the debut set from 1989 and the iconic Ken Griffey rookie card):

Yet Pinnacle topped them:

Fleer Ultra was better still:

And 1994 Leaf was a work of art:
I contend the base sets were even better designed (for what they were) in 1994. Since I didn’t have a whole lot of disposable income, almost all of my pack purchases were of these five sets. Comparing the number of each I appear to have in my collection, I can roughly assume that my order in preference working from least-favorite to favorite was:
5. Donruss

This one was (and remains) a distant fifth on my order of preference, even though I think it’s a very nice-looking card. I have roughly equal amounts of the next three:
4. Upper Deck Collector’s Choice

I think it’s the pinstripes that really makes these pop. Plus the old-timey drawing to go along with the player’s position.
3. Fleer

The only flaw is that the player’s name is understated by being wrapped in small letters around the team’s logo.
2. Score

Sadly, what made this set great is what makes them less desirable to collect twenty-six years later. I was not in anyway part of the “cards in the bike spokes” generation. I treated my cards, particularly the cards of stars, as if they were my most valuable possession (in fact, they probably WERE my most valuable possession). And yet I’m not sure I came across one in my album that didn’t have obvious chipping to those amazing dark borders.
As you’ve probably guessed from the title of this post, Topps is #1.
After 1994, it was all downhill for baseball cards, at least for the rest of the nineties when I was still interested in them although no longer obsessed. What went wrong? The most personal is that I became much more interested in the numbers on the back of the cards than in the cards themselves. But the biggest problem is that the crisp designs of 1994, which focused on the player pictures and used borders, names, and logos to complement them, were benched in favor of drawing attention to everything but the picture. Perhaps the designers wanted to show off what they could do in MS Paint (and some really do look like they were designed in MS Paint)? Or maybe they decided that since everyone had nice pictures on their cards, it was necessary to seek a graphic design that would differentiate them from the pack (pun intended)?
Perhaps we should have seen it coming, as Topps juxtaposed their beautiful 1994 base set with the questionable Stadium Club and the horrifying Topps Finest:
When looking at 1995, I think it’s instructive to look at what happened to the base sets, which were all so great in 1994. In 1995, worst went to first by default – Donruss changed it up a little bit, while their competitors decided to jump off a cliff together:

If I remember correctly, before the strike was settled, Upper Deck leaped first in the spring with a “special edition” of Collector’s Choice. Out are the classic pinstripes; in is garish blue. In is haphazard capitalization. Out were my dollars:

Score may have realized that the black borders were a disaster for the long-term condition of their cards; I’m not sure why that required shrinking the pictures, adding a faux wood/dark green border, and circles of varying sizes for some unknown reason. What’s sad here is that they were so close to some classic baseball motifs – wood grain can work (see 1987 Topps), but it helps if it looks like what you’d see on a bat. Green is the color I most associate with baseball – but the green of the grass, not a pine tree.

People didn’t say “hold my beer” in 1995, but if they did, Topps would have, going from the most perfect set ever printed to a terrible font in gold (so much for the first and classiest parallel set, Topps Gold), often hard to read because it brushes up against the border which for some reason is not straight:

Still, nothing better captures the 1995 self-own of the big five than the monstrosity that was 1995 Fleer. 1994 Fleer, as I said above, was gorgeous but almost too simple. They fixed that right quick. In 1995, Fleer decided that they would obscure the front of the card with all of the biographical info that no one cared about (and often didn’t believe). But that wasn’t enough – they decided that each of the six divisions should have a unique design. None got it worse than the AL Central, which was not good news for the cards of my Indians heroes:

Perhaps all of the card designers were on strike in solidarity with the players? I think 1995 was a low point – it got a little better later in the decade. But even the bible of the hobby lost its way. When I was taken by cards, I naturally asked for a subscription to Beckett Baseball Card Monthly. I’ll always remember the cover of the first issue I bought separately – a great portrait of Jeff Bagwell. These frameable covers would persist beyond 1994, but not too much longer, and eventually even Beckett covers would have headlines everywhere, like 1995 Fleer had grown beyond the borders of its set and conquered all things baseball card.

Tuesday, July 28, 2020
Chutzpah
Step 1: Advocate tirelessly for baseball to be shutdown, along with anything else that your political masters deem “non-essential” (all while being completely oblivious to how totalitarian this all is, as you not only no longer claim to be guided by liberal values, ideals, or principles, you and the political movement you follow have completely lost the ability to even think in terms of liberal values, ideals, or principles).
Step 2: Baseball (and other “non-essential” means of voluntary economic cooperation between individuals that provide the livelihoods for the people who buy your subscriptions and advertise on your website) gets shutdown.
Step 3: Shockingly, your revenues from subscriptions and advertisements declines.
Step 4: Ask me for money so that you can cover the activity that you tirelessly advocated to be shutdown.
Hard pass. I’d wish you good luck, but I wouldn’t mean it.
Thursday, July 23, 2020
2020 Predictions
I really should eschew doing predictions this year – the whole point of an exercise like this (other than fun, which is the main point) is to predict what will happen over a reasonably large sample. I don’t predict the outcomes of playoff series in the same manner, because I contend they are inherently unpredictable as binary outcomes with any level of accuracy that makes it worthwhile. The practice of making rank-order predictions is already a simplification of the reality of what is actually being predicted, and when applied to a season that is slated to be less than 40% the normal length, it is a foolhardy exercise indeed. Add on extra uncertainty due to player availability variability beyond the normal injuries, an extended gap since the last time we actually had the opportunity to observe players’ talent levels on the field, a severely unbalanced schedule, etc. ad nauseam, and there’s no good reason to do it.
Except that it’s fun, and I’ve been doing it in an unbroken chain since 1995, and if there's ever been a season in which to try to embrace the fun elements of baseball, this is it. So why not? I didn’t put a lot of my own effort into this – usually I use the Marcel or ZIPS or Steamer projections as a starting point, but make my own tweaks to both player’s performance and my thoughts on likely playing time. Here I just used the team win estimates published by Baseball Prospectus, Fangraphs, and Clay Davenport as a starting point rather than building up from player-level performance. I also made some of my own judgment calls on team-level performance more aggressively than I normally would – it’s easier to disbelieve someone’s team-level prediction when you haven’t dug in at the player-level yourself. I have not in any way though inserted randomness for what I hope would be obvious reasons if you are reading this blog.
AL EAST
1. Tampa Bay
2. New York (wildcard)
3. Toronto
4. Boston
5. Baltimore
AL CENTRAL
1. Minnesota
2. Cleveland (wildcard)
3. Chicago
4. Kansas City
5. Detroit
AL WEST
1. Oakland
2. Houston
3. Los Angeles
4. Texas
5. Seattle
NL EAST
1. New York
2. Atlanta (wildcard)
3. Washington
4. Philadelphia
5. Miami
NL CENTRAL
1. Chicago
2. Cincinnati (wildcard)
3. Milwaukee
4. St. Louis
5. Pittsburgh
NL WEST
1. Los Angeles
2. Arizona
3. San Diego
4. Colorado
5. San Francisco
WORLD SERIES
Los Angeles over Tampa Bay
Wednesday, July 15, 2020
Hallmarks of Quality Metrics
Another old post I never published, probably because it was repetitive of sentiments I'd written before. I'm guessing I must have encountered a metric that really annoyed me and this was written as a responsive missive.
In a previous article I discussed some of the shortcomings of OPS as an advanced metric, which naturally leads to the question: “What are the characteristics of good advanced metrics?” While the relative importance that one places on each criterion is up for debate (the list that follows is in no particular order), the following considerations should be relatively non-controversial. I’ve used the term “metric” to refer to any statistic or derived category, which is not precise terminology:
1. Clear purpose
Before one designs a metric or uses it to answer a question, it’s imperative that the question of interest be defined. What is the metric setting out to measure? Most metrics in use, even those that are not in favor with sabermetricians, do fairly well on this score. Counting statistics, regardless of their ultimate utility, are largely clear in terms of definition and meaning. Some, like hits or strikeouts, are inherently obvious. Those with more involved definitions often still have a clear purpose even if the execution of that idea is somewhat muddled (like errors).
2. Developed with a theory in mind
This criterion is closely related to a clear purpose, but takes it a step further by questioning the thought process that went into developing the metric. OPS doesn’t fail, as it is based on the reasonable notion that hitting can be broken down into the broad categories of getting runners on base (OBA) and advancing them (SLG). However, due to the somewhat arbitrary nature in which the two statistics are combined, OPS does not match up to metrics like wOBA and True Average which have as their basis a linear weight model of the run scoring process. Some proposed metrics fail spectacularly, though, as they simply combine statistical columns without any particular rhyme or reason. Thankfully, most of these fail to gain traction, but some fail to gain traction yet still have their own Wikipedia pages. Metrics of this type may appear to “work” as they will generally produce reasonable leader boards, but the same could be said for any haphazard combination of positive events and categories.
3. Accurate
A metric should result in an accurate estimate of whatever it is designed to measure. For instance, a metric that attempts to measure offense productivity should have a strong correlation with team runs scored as scoring runs is the prime objective for an offense. The best-performing models for estimating team runs scored tend to be based on either dynamic models of the run scoring process (such as David Smyth’s Base Runs) or linear weight models (pioneered by George Lindsey and Pete Palmer and now in wide use). Thus it stands to reason that metrics built on linear weights (such as wOBA) are a better tool to use when evaluating offensive production than alternatives that do not correlate as well with runs scored.
Sometimes, though, it is not easy to measure accuracy due to a lack of data to verify against or a desire to use the metric to address a similar but subtly distinct question. For example, metrics validated against team results are often used to measure individual performance, which leads to the next criterion.
4. Adaptable over a wide range of contexts
While there is nothing inherently wrong with a metric that is designed to work only under a limited set of conditions--so long as said metric is not stretched beyond its capabilities--it is better still to be confident that the metric will produce reasonable results for a broader set of questions.
Sometimes metrics work well over normal ranges of performance and thus provide reasonable answers for most questions. For example, the common rule of thumb that 10 runs = 1 win is quite accurate at predicting the win totals of major league teams from their runs scored and allowed. However, the actual relationship between runs and wins is not linear—it only appears to be linear because the conversion is calibrated over a narrow set of possible outcomes. When the model is applied to more extreme conditions (which in this case could be an average level of runs scored per game much different than major league norms or teams with very low or very high run differentials), the accuracy will suffer. A dynamic model of estimated winning percentage (such as Pythagenport) can maintain accuracy over a wider range of scenarios.
A related but slightly different issue occurs when some metrics that are designed for use with team data are applied to individuals. A classic example is Bill James’ original version(s) of Runs Created, which recognizes the dynamic relationship between getting runners on base and advancing them. When applied to an individual’s statistics, though, the implication is that the player is reaching base, then advancing himself around the bases, whereas he actually interacts with his teammates. The resulting distortion requires that caution be used when interpreting Runs Created estimates for individual players.
5. Expressed in meaningful units
Ideally, the metric should return a result that has a logical, interpretable baseball meaning. Metrics expressed in terms of runs and wins are ideal since the connection to the objective of the game is made clear, but there are any number of other expressions that can be meaningful. On Base Average, for instance, represents the percentage of plate appearances in which a batter reaches safely, which is easy to explain and easy to think about in terms of on-field implications.
In some rare instances, it is next to impossible to express a result in meaningful units and so a nebulous value must suffice. One example is Bill James’ Speed Score, which endeavors to estimate a player’s speed skill by taking into account a number of categories related to speed (such as stolen base attempt frequency, rate of triples per ball in play, defensive range, etc.) Since there is no single manifestation of speed on the field and no obvious units to capture baseball speed, James uses an abstract scale.
6. Not needlessly complex
It is certainly tempting to say that metrics should be simple, but in my opinion simplicity need not be a goal unto itself. What is important is that the metric not make things more complicated than they need to be.
However, describing complex processes sometimes necessitates the use of complex models. The key is to avoid complexity for its own sake and phony precision. The end use and user of the metric should also be considered—if a “quick and dirty” estimate will suffice, then a simple metric may suffice, but a more complex metric can be used when a true best estimate is needed.
7. Catchy Name
This final entry is somewhat tongue-in-cheek, as it is irrelevant to the quality of a metric, but there’s no denying that when it comes to mainstream acceptance, marketing matters. To bring things full circle, a good name succinctly references the intended purpose and use of the metric while providing a minimum amount of ammunition to those looking to mock the field. Whether any sabermetric measures score particularly well on this front will be left as a rhetorical question for the reader.
Wednesday, July 08, 2020
April 4, 1994 pt. 2
I’ve previously written about the Indians/Mariners opening day game of April 4, 1994 that made me a baseball fan. I won’t rehash all that in detail again, but I recently was able to watch a replay of this game for the first time. I’d never actually seen any of it before, except for highlights – in real time I listened to about the seventh inning forward on the radio.
For the rewatch, I kept a scoresheet, which is reproduced at the bottom of the post. A few observations:
* Chris Berman and Buck Martinez called the game on ESPN. Berman was not as terrible as I remember him being, but most of my exposure was later. That is not to say that he was good. Martinez is a middling announcer with a terrible voice, and was in 1994 as well. It would be a real treat to be able to watch this game with the local radio call of Herb Score and Tom Hamilton that I would have enjoyed in place of the national guys.
* Randy Johnson had a no-hitter through seven, which was noteworthy for reasons beyond the obvious. As all Indians fans know, Bob Feller is the only pitcher to throw an Opening Day no-hitter, and here was a threat to no-hit the Indians in on Opening Day in their first game in their new park with Feller on hand. Plus Randy Johnson, while not yet the legend that he would be, was obviously a legitimate no-hit candidate. He’d already thrown one in 1990, and 1993 had been his breakout year, finishing second in the Cy Young voting and recording his third straight season with over 10 K/9.
So not having seen the game and filling in the details in my mind given what I knew about the Big Unit later, I assumed that he had spent the first seven innings carving Cleveland up. But that was not the case at all; it more resembled what you would have expected a Greg Hibbard no-hit bid to look like. Through seven, Johnson had walked four and fanned two on 94 pitches. His twenty-one outs were distributed as:
12 on groundouts (including 2 DPs)
5 on flyouts
2 on strikeouts
1 popout
1 caught stealing
His opposite number, Dennis Martinez, was pitching a similar game from a DIPS perspective with one big exception – the two out solo shot he yielded to Eric Anthony in the third. Otherwise, through seven Martinez had struck out four, walked four, and hit Edgar Martinez in the first inning (providing an early injury scare as Mike Blowers pinch-ran, all this after Martinez had appeared in just 42 games in 1993. He’d only appear in three more games the rest of April).
* Two future stars were languishing down in the Indians lineup – Manny Ramirez batting eighth, and Jim Thome on the bench. It would be some time before Thome was trusted to start against left-handed pitchers, and so Mark Lewis was the ninth-place hitter and third baseman. Ramirez provided a Manny being Manny moment. After Candy Maldonado walked to open the eighth and Sandy Alomar singled to break up the no-no, Manny clanged a 1-0 Johnson offering off the big wall in left for a game-tying double. With Mark Lewis looking to advance the go-ahead run to third base, Ramirez strayed two far off second and was picked off by a Dan Wilson throwback to second on the first pitch.
Ramirez and Thome were never in the game simultaneously; with the Indians down a run with one out in the tenth, Ramirez drew a walk and was replaced by pinch-runner Wayne Kirby. It was then that Thome batted for Lewis, which brought on lefty reliever King for Seattle. Thome pulled a double down the right field line to put runners at second and third, and Kirby would later score when Vizquel hit into a fielder’s choice. It would work out in the end two, as Kirby walked it off in the eleventh with a two-out, line drive single to left to score Eddie Murray with the winning run.
* Despite what I’m about to say below, this was a good game for star power as these two teams would emerge as top AL contenders of the latter half of the nineties: Hall of Famers Randy Johnson, Ken Griffey, Edgar Martinez, Eddie Murray, Jim Thome, future Hall of Famer Omar Vizquel, would have been Hall of Famer Manny Ramirez, should be Hall of Famer Kenny Lofton, could have been Hall of Famer Albert Belle, a former Rookie of the Year in Sandy Alomar, and other memorable names including Jay Buhner, Carlos Baerga, Tino Martinez, and Jose Mesa.
* One thing that struck me in re-watching it is what an ordinary game it was. Granted, given the circumstances (opening day and opening game of a new park) it was extremely memorable for Indians fans, but if you strip all that out and just evaluate it as a game, it wouldn’t be the most exciting of most major league team’s seasons. I have personally attended at least six Indians games in the last four seasons that were more compelling, and I’ve only been to about sixty games in that time and I’m making that list from the top of my head. I had built it up in my head as a kind of epic, and in some senses it disappointed upon rewatch.
On the other hand, that disappointment reminded me of what a great game baseball is. I have now watched twenty-six seasons of major league baseball and perhaps become jaded about just how interesting and exciting baseball inherently is. That this game wouldn’t rank in the top 10% of games I’ve attended recently speaks to what an amazing game baseball is. Since this game was sufficient to almost instantly turn me into a baseball nut, I suspect that a much less exciting contest would have done the trick. And it should have...I’m repeating myself again, and as I write this we are still two and a half weeks from even the possibility of baseball in 2020, and that too reminds me that baseball is just the best in every way.
* More generally on the franchise that I yolked myself to on April 4, 1994, I have no comment on the fact that the Indians will liekely soon be changing their name itself. I do have two strains of thought on possible future names:
1. I think “Expos” is the logical choice, which is a snarky way of saying that my suspicion is that this name will be changing again in the relatively near future as the franchise settles into its new home in Montreal, Nashville, Portland, Las Vegas, Charlotte, etc.
2. “Spiders” is a dreadful option. First of all, as a general philosophy, I believe that baseball team names should be non-threatening. Most baseball team names are – I would contend that the only exceptions among the sixteen teams dating to 1901 or earlier are Pirates and Tigers, depending on what you think (very carefully) about Braves and Indians. Cubs are not an animal I would wish to encounter, but the name suggests cute and cuddly teddy bears rather than miniature grizzlies. Among expansion team names, the only one that I would classify as even mildly threatening is Rangers, and I would suspect the desired effect is strength and honor rather than menace.
The exception is the 1998 expansion. The Devil Rays and the Diamondbacks both sound threatening, although the former is actually generally harmless (to humans at least, and I think that’s all we should consider lest all the bird names become threatening) and was later downgraded to the double meaning “Rays” anyway. The latter is a scary animal, but is also in my opinion a contender for best expansion team name, due to the baseball tie in (my other contender for best expansion team names would be Brewers (although that was recycled), Colt .45s/Astros, and Pilots/Mariners. The latter was a case in which the city had a great name and then got a similar yet superior one eight years later).
So I would contend that Spiders is contrary to the spirit of baseball nicknames. The history of the name is also quite problematic (although quite appropriate if my misgivings about the future of the franchise are founded). The original Spiders represented Cleveland in the National League from 1887-1899, never winning a pennant. In the early 1890s they were a strong outfit, finishing second three times and even capturing a Temple Cup (which I do not in any way deem to be comparable to a regular season pennant) with names like Cy Young and Jesse Burkett, but were soon a victim of the systemic corruption of the 1890s NL, with owner Stanley Robison siphoning off talent for the St. Louis now-Cardinals in which he also had a stake. As you probably know, this culminated in the 20-134 debacle of 1899 before the team joined Detroit, Lousiville, and Washington on the chopping block, leaving Cleveland open for Ban Johnson’s play at major status for the American League two years later. I would contend that this is quite an ignominious history and nothing to be celebrated or emulated.
If Cleveland’s major league history must be the first source of inspiration, the Indians’ prior unofficial names won’t cut it: Blues is boring, Cleveland isn’t supporting a team called the Broncos, Naps would be fine with me but doesn’t sell and the headlines write themselves. The Players League outfit was referred to as the Infants. The Negro Leagues don’t provide much in the way of an option, as Cleveland’s proudest entry was the Buckeyes, a name of which THE sports team of only entity is worthy.
I do think there is one Cleveland major league name that would work – the first, the Forest City club which represented the city in the National Association during 1871-72. This team actually participated in the NA’s first league game on May 4, 1871. Maybe you’d have to rework it to Foresters (or even Sawyers), but it’s a name I could get behind.
Best non-historical choice, although semi-violating my own suggested rule about nice namesakes: Buzzards.


Wednesday, July 01, 2020
"Replacement Level" Managers
This is an old post that I never published. It's not good, as it just presents something of a freak show stat, but I was mildly interested by it when I re-read it so maybe someone out there will be as well. All of the facts/figures are through 2009 and I did not update them at all. I did not one factual error which is also not corrected - Billy Southworth was inducted into the HOF in 2008.
I put quotes around "replacement level" in the title because this article is not really about establishing a replacement level for managers in the same sense as the phrase would imply when discussing players. It is rather about establishing a baseline for crude comparisons of managerial records, in the same vein as WAR--but without any claim that the baseline represents the point at which talent is freely available.
After all, it's folly to hold up a manger's W-L record as the sole evidence of his quality as a manager. Even the most ardent believers in the importance of managers to a team's record cannot possibly believe that they can separate the manager's contribution from all of the other noise that goes into a team's record.
If you want a crude method to compare managerial W-L records, there are few options that come to mind. Conventional approaches would include just looking at total wins, winning percentage, and games over .500, just as one might do with pitcher W-L records.
Of course, my own initial thought as a sabermetrician is to turn to a baseline that values longevity to some extent. If a manager is allowed to direct 3,942 major league games, yet has a sub-.500 record, it would be silly to assign him a negative number and move on (Gene Mauch). Managers are obviously employable even with losing records, and there are many factors well outside the manager's control that contribute to a team's record.
So my natural inclination is to look at a manager's wins above replacement, which inevitably leads to a decision about how to define managerial replacement level. There are a lot of ways to estimate replacement level for players, but one of the simplest is to look at the aggregate performance of players given very little playing time. The analogous solution would be to look at managerial records for those managers that were replacements, managing less than a full season of games.
When using this approach for players, one must be careful to consider the selective sampling issues involved--players that fail in an initial trial are less likely to receive future playing time, even though it is possible that their true talent is greater (the opposite is also true to some extent). The same is also likely true to some extent for managers--managers whose teams do not perform well in an initial interim role are not as likely to be retained. However, since my application here is just establishing a rough baseline to use for ultimately unimportant comparisons of managerial records, I am simply going to proceed as if these concerns are irrelevant.
The goal is not to devise a rating system for managers; it is to find a crude baseline to use for comparing un-contextualized managerial records. The freak show nature of the exercise is evident, and hopefully will serve to excuse my playing fast and loose with proper research procedure.
What I did was look at career records for all managers with less than 154 games managed (Although I then removed managers who served full season stints in seasons with less than 154 games from the list as well, as well as Cubs managers from the early 60s who were part of the College of Coaches experiment and Stanley Robison and Ted Turner, who owned their teams and weren't real managers.) from 1901-2009. This is my group of "replacement-level" managers. There are 109 such managers, serving in a total of 135 different team-seasons. Their career totals of games managed range from one (ten managers, with either Rudy York or Eddie Yost as the biggest name) to 149 (Tom Runnells with the 1991-92 Expos).
Overall, they managed 5530 games (an average of 41 games each), going 2322-3208 for a .420 W%. So that will be my baseline for managerial records--.420.
By using .420 as a baseline, I don't mean to imply that it is a replacement-level in the traditional sense. It is quite possible that interim managers generally don't keep their jobs if they don't manage at least a .420 W%, but I don't mean to imply that replacement managers are ".420 managers".
If one was to attempt to measure a manager's replacement level in terms of actual effect on a team attributable to the skipper, my intuition is that it would be close to .500. There are simply too many possible candidates for managerial positions for me to think otherwise. Regardless, though, this "study" in no way indicates that the managers lowered .500 teams to .420.
The teams had a total aggregate record (with both the replacement and non-replacement managers) of 9334-11752, a .443 W%. This comparison does not take into account that the games managed by replacements ranged from one to over 140.
A crude way to compare team performance with and without the replacement level manager is to weight each team-season by the minimum of games managed by the replacement and other games. Using this approach, the weighted average of (W% with replacement manager - W% otherwise) is -.019.
Another crude approach is to weight by the harmonic mean of games managed by the replacement and others, rather than the minimum of the two. The weighted average difference is -.025 when using the harmonic mean. Those results should not be used to draw any conclusions, but without any regression or significance testing they imply that a replacement-level manager might lower a .500 team to .480 or .475, a difference in the range of four games a year. I am not claiming that is true, for the selective sampling reasons discussed previously among a myriad of other reasons.
With that out of the way, I will present some data on managerial records above .420 for managers, 1901-2009. I'll call this Austin Rating in honor of Jimmy Austin, who is the only man to serve three such stints as manager (all with the Browns) without reaching 154 career games. Austin's player-manager career started with St. Louis in 1913, replacing George Stovall temporarily (2-6) before Branch Rickey took over permanently. He also did a stint in 1918 (7-9) in relief of Fielder Jones before Jimmy Burke stepped in. His final and longest experience at the helm was in 1923, when he was 22-29 replacing Lee Fohl. His career 31-44 mark (.413) is a little below the .420 baseline, so his own Austin Rating is -.5.
Here are the top 25 career managers (again, through 2009):
There are sixteen Hall of Fame managers from this period; fourteen are in the top 25 for Austin Rating, with Whitey Herzog (270, 28th) and Wilbert Robinson (224, 34th) just missing the top 25. This is not offered as an indication that Austin Rating tracks HOF managerial choices or that it correctly identifies good managers, as any reasonable system based on career wins and losses would likely produce similar results for Hall of Fame skippers.
Going down the list, the non-Hall of Famers are either active or recently retired (Cox, LaRussa, Torre, Piniella) or in the Hall of Fame as a player (Clarke) until you get to Billy Southworth (Clark Griffith is also in the Hall, with a noteworthy career in the areas of playing, managing, and ownership). Southworth does not have wins in bulk (which seem to be the true indicator of HOF selection), but his .597 W% results in a very strong Austin Rating.
Here are the bottom ten managers:
Most of these guys served in the early part of the twentieth century, when competitive balance was less pronounced and multiple franchises had long walks in the wilderness. Protho brings up the rear for managing three teams in Phillies dreadful pre-War stretch (1939-41). The only manager on the list that commanded over half of his games post-1950 was Roy Hartsfield, original skipper of the expansion Blue Jays. Extending the list down to 13th would include Alan Trammell, while Manny Acta ranks 18th lowest, but including 2010 would give him a slight bump as the Indians scraped over the .420 mark.
Finally, here is the leader in Austin Rating for each current team in their current city (except Washington which doesn't have much of a history; record with that franchise only):
Tuesday, June 16, 2020
Preoccupied With 1985: Linear Weights and the Historical Abstract
I stumbled across this unpublished post while cleaning up some files – it was not particularly timely when written about ten years ago, and is even less timely now. Unlike some other old pieces I find, though, I don’t know why I never published it, other than maybe redundancy and beating a dead horse. I still agree with the opinions I expressed, and it is well above the low bar required for inclusion on this blog.
The original edition of Bill James’ Historical Baseball Abstract, published in 1985, is my favorite baseball book, and I am far from the only well-read baseball aficionado who holds it in such high regard. It contains a very engaging walk through each decade in major league history, some interesting material on rating players (including what has to be one of the first explicit discussions of peak versus career value in those terms), ratings of the best players by position and the top 100 players overall, and career statistics for about 200 all-time greats which seem like nothing in the internet age but at the time represented the most comprehensive collation on those players.
However, there is one section of the book which does not hold up well at all. It really didn’t hold up at the time, but I wasn’t in a position to judge that. James reviews The Hidden Game of Baseball, published the previous year by John Thorn and Pete Palmer, and gives his thoughts about the Linear Weights system.
James’ lifelong aversion to linear weights is somewhat legendary among those of us who delve deeply into these issues, but the discussion in the Historical Abstract is the source of the river, at least in terms of James’ published material. For years, James’ thoughts colored the perception of linear weights by many consumers of sabermetric research. This is no longer the case, as many people interested in sabermetrics twenty-five years later have never read the original book, and linear weights have been rehabilitated and widely accepted through the work of Mitchel Lichtman, Tom Tango, and now many others.
So to go back thirty years later and rake James’ essay over the coals is admittedly unfair. You may choose to look at this as gratuitous James-bashing if you please; that is not my intent, but I won’t protest any further than this paragraph. I think that some of the arguments James advances against linear weights are still heard today in different words, and occasionally you will still see a reference to the article from an old Runs Created diehard. And if one can address the concerns of the Bill James of 1985 on linear weights, it should go a long way in addressing the concerns of other critics.
It should be noted that James on the whole is quite complementary of The Hidden Game and its authors. I will be focusing on his critical comments on methodology, and so any excerpts I use will be of the argumentative variety and if taken without the disclaimer could give the wrong impression of James’ view of the work as a whole.
The first substantive argument that James offers against Palmer’s linear weights (in this case, really, the discussion is focused on the Batting Runs component) is their accuracy. The formula in question is:
BR = .46S + .80D + 1.02T + 1.40HR + .33(W + HB) + .3SB - .6CS - .25(AB - H) - .5(OOB)
As you know, Palmer’s formula uses an out value that returns an estimate of runs above average rather than absolute runs scored (in which case it would be somewhere around -.1). The formula listed by Palmer fixes the out value at -.25, but it is explained that the actual value is to be calculated for each league-season. James notes this, but then ignores it in using the Batting Runs formula to estimate team runs scored. To do so, he simply adds the above result to the league average of runs scored per team for the season. He opines that the resulting estimates are “[do] not, in fact, meet any reasonable standard of accuracy as a predictor of runs scored.”
And it’s true--they don’t. This is not because the BR formula does not work, but rather because James applied it incorrectly. As he explains, “For the sake of clarity, the formula as it appears above yields the number of runs that the teams should be above or below the league average; when you add in the league average, as I did here, you should get the number of runs that they score.”
This seems reasonable enough, but in fact it is an incorrect application of the formula. The correct way to use a linear weights above average formula to estimate total runs scored is to add the result to the league average runs/out multiplied by the number of outs the team actually made.
This can be demonstrated pretty simply by using the same league-seasons (1983, both leagues) that James uses in the initial test in the Historical Abstract. If you use the BR formula using -.25 as the out weight and simply add the result to the league average runs scored (in each respective league), the RMSE is 29.5. Refine that a little bit by adding in the number of outs each team made multiplied by the respective league runs/out (but still using -.25 as the out weight), the RMSE improves to 29.3. The James formula that uses the most comparable input, stolen base RC, has a RMSE of 24.4, and you can see why (in this limited sample; I’m certainly not advocating paying much heed to accuracy tests based on one year of data, and neither was James) he thought BR was less accurate. But had he applied the formula properly, by figuring custom out values for each league (-.255 in the AL and -.244 in the NL) and adding the resulting RAA estimate to league runs/out times team outs, he would have gotten a RMSE of 18.7.
In fairness to James, the authors of The Hidden Game did not do a great job in explaining the intricacies of linear weight calculations. The book is largely non-technical, and nitty-gritty details are glossed over. The proper method to compute total runs scored from the RAA estimate is never exactly explained, nor is the precise way to calculate the out value specific to a league-season (while it’s a matter of simple algebra, presenting the formula explicitly would have cleared up some confusion). To do a fair accuracy test versus a method like Runs Created, which does not take into account any data on league averages, you would also need to calculate the -.1 out value over a large sample and hold it constant, which Thorn and Palmer did not do or explain. In addition, the accuracy test was not as well-designed as it could have been, although that wouldn’t have had much of an impact on the results for Batting Runs or Runs Created, but rather for rate stats converted to runs.
James then goes on to explain the advantage that Batting Runs has in terms of being able to hone in on the correct value for runs scored, since it is defined to be correct on the league level. He is absolutely correct (as discussed in the preceding paragraph) that this is an unfair advantage to bestow in a run estimator accuracy test; however, it is also demonstrable that even under a fair test, Batting Runs and other similar linear weight methods acquit themselves nicely and are more accurate than comparable contemporary versions of Runs Created.
In the course of this discussion, James writes “What I would say, of course, is that while baseball changes, it changes very slowly over a long period of time; the value of an out in the American League in 1987 will be virtually identical with the value of an out in the American League in 1988.” This turned out to be an unfortunate future example for James since the AL averaged 4.90 runs/game in 1987 but just 4.36 in 1988. James’ point has merit--values should not jump around wildly for no reason other than the need to minimize RMSE--but the Batting Runs out value does not generally behave in a matter inconsistent with simply tracking changes in league scoring.
James’ big conclusion on linear weights is: “I think that the system of evaluation by linear weights is not at all accurate to begin with, does not become any more accurate with the substitution of figures derived from one season’s worth of data…Linear weights cannot possibly evaluate offense for the simplest of reasons: Offense is not linear.”
He continues “The creation of runs is not a linear activity, in which each element of the offense has a given weight regardless of the situation, but rather a geometric activity, in which the value of each element is dependent on the other elements.” James is correct that offense is not linear and that the value of any given event is dependent on the frequency of other events. But his conclusion that linear weights are incapable of evaluating offense is only supported by his faulty interpretation of the accuracy of Batting Runs. While offense is not linear, team offense is restricted to a narrow enough range that linear methods can accurately estimate team runs scored.
More importantly, James fails to recognize that while offense is dynamic, a poor dynamic estimator (such as his own Runs Created) is not necessarily (and in fact, is not) going to perform better than a linear weight method at the task of estimating runs scored. He also does not consider the problems that might be inherent in applying a dynamic run estimator directly to an individual player’s batting line, when the player is in fact a member of a team rather than his own team. Eventually, he would come to this realization and begin using a theoretical team version of Runs Created (which is one of the many reasons this criticism of his thirty-five year old essay can be viewed as unfair).
Much of the misunderstanding probably could have been avoided had Batting Runs been presented as absolute runs rather than runs above average. Palmer has never used an absolute version in any of his books, but of course many others have used absolute linear weight methods. One of the more prominent is Paul Johnson’s Estimated Runs Produced, which was brought to the public eye when none other than Bill James published Johnson's article in the 1985 Abstract annual.
Johnson’s ERP formula was dressed up in a way that made it plain to see that it was linear, but did not explicitly show the coefficient for each event as Batting Runs did. Still, it remains almost inexplicable that an analyst of James’ caliber did not see the connection between the two approaches, as he was writing two very different opinions on the merits of each nearly simultaneously.
James also applies his broad brush to Palmer’s win estimation method, saying that if you ask the Pythagorean method “If a team scores 800 runs and allows 600, how many games will they win?”, it gives you an answer (104), while “the linear weights” says “ask me after the season is over.”
The use of the phrase “wait until the season is over” is the kind of ill-conceived rhetoric that seems out of place in a James work but would be expected in a criticism of him by a clueless sportswriter. Any metric that compares to a baseline or includes anything other than the player’s own performance (such as a league average or a park factor) is going to see its output change as that independent input changes. That goes for many of James’ metrics as well (OW% for instance).
To the extent that the criticism has any validity, it should be used in the context of Batting Runs, since admittedly Palmer did not explain how to use linear weights to figure an absolute estimate of runs in the nature of Runs Created. To apply it to Palmer’s win estimator (RPW = 10*sqrt(runs per inning by both teams)) simply does not make sense. The win estimator does not rely on the league average; it accounts for the fact that each run is less valuable to a win as the total number of runs scored increases, but it doesn’t require the use of anything other than the actual statistics of the team and its opponents. (Of course, when applied to an individual player’s Batting Runs it does use the league average, which again is no different conceptually than many of James’ methods.) The Pythagorean formula with a fixed exponent has the benefit (compared to a linear estimator, even a dynamic one) of restricting W% to the range [0, 1], but it also treats all equal run ratios as translating to equal win ratios.
James concludes his essay by comparing the offensive production of Luke Easter in 1950 and Jimmy Wynn in 1968. His methods show Easter creating 94 runs making 402 outs and Wynn creating 91 runs making 413 outs, while Batting Runs shows Easter as +29 runs and Wynn +26.
James goes on to point out that the league Easter played in averaged 5.04 runs per game, while Wynn’s league averaged 3.43, and thus Wynn was the far superior offensive player, by a margin of +37 to +18 runs using RC. “Same problem--the linear weights method does not adapt to the needs of the analysis, and thus does not produce an accurate portrayal of the subject.”
In this case, James simply missed the disclaimer that the out weight varies with each league-season. While it makes sense to criticize the treatment of the league average as a known in testing the accuracy of a run estimator, it doesn’t make any sense at all to criticize using it when putting a batter’s season into context. Of course, James agrees that context is important, as he converts Easter and Wynn’s RC into baselined metrics in the same discussion.
When Batting Runs is allowed to calculate its out value as intended, it produces a similar verdict on the value of Easter and Wynn. In Total Baseball (using a slightly different but very much same in spirit Batting Runs formula), Palmer estimates Wynn at +38 and Easter at +14, essentially in agreement with from James’ estimate of +37 and +18. The concept of linear weights did not fail; James’ comprehension of it did. It doesn’t matter if that happened because Palmer and Thorn’s explanation wasn’t straightforward (or comprehensive) enough, or whether James just missed the boat, or a combination of both. Whatever the reason, the essay “Finding the Hidden Game, pt. 3” is not a fair or accurate assessment of the utility of linear weight methods and stands as the only real blemish on as good of a baseball book as has ever been written.