From time to time you will hear about the baseball gods, who apparently control luck, show favor upon some teams and not upon others, and other such godly tasks. As playoff time approaches and fans everywhere seek favor for their teams with the gods, the unanswered question is “Who are these gods anyway?”
Luckily, I have uncovered a new ancient text that sheds light on this important issue. Why one of the baseball gods bestowed upon me the honor of finding this information I cannot say. I know little about mythology; I did take Classics 224, but that was titled something to the effect of “ancient Greek civilization”; the gods were discussed, but not exclusively. Perhaps if I would have taken the mythology-focused course, Classics 222 (better known as “who screwed who 222”), I would be better suited to this responsibility. Hopefully what follows would not make Professor Tran cry.
Nonetheless, it is my responsibility to share the identities of the gods with the world. Each god had a counterpart in Greek mythology, with whom they share at least one characteristic, even if it is a bit of a stretch.
The father of the gods is Zeus. Who was the father of baseball? There are a couple of people who have been bestowed with this title, but there is one to whom it is more universally applied. I do not believe he had the ability to shoot thunderbolts at his enemies, and I am unaware of any tales of his extramarital affairs. Nonetheless, the father is the father. Zeus is Henry Chadwick.
When Zeus overthrew his father Cronus and the Titans, it brought about a new era of order and prosperity. The old order was gone, and the threats to the new gods was gone. Likewise, when Father Chadwick and company established baseball as our national game, it became the unquestioned king of the ball and stick games. Thus, Cronus is town ball, the Massachusetts game, cricket, and other pretenders to the throne.
When Cronus castrated his father, out of the bloody mess emerged a goddess, a goddess of love. Our baseball god is a man, yet he served a similar function as one of the first sex symbols of the new national game. Some even credit him with the creation of Ladies’ Days, although this is probably not true. Nonetheless, he was a good pitcher, although not great enough that we don’t remember him too much for his pitching. He even carried the nickname “The Apollo of the Box”. Aphrodite is Tony Mullane.
There was another god who in some tellings was destined to be even greater than Zeus. His followers were often under a trance, worshiping him and the wine and pleasure that came with him. In baseball, there would arise a player so great that some would claim that he and not Father Chadwick should occupy the highest throne on Olympus. Alas, it was not to be, but his many followers and his indulgence in pleasure and sin were as prodigious as his home runs. Dionysus is Babe Ruth.
War has been a reality since the dawn of the gods and of civilization, and of course baseball like any team sport can be viewed as a bloodless, refined, civilized proxy for war. There was a god who loved war and bloodlust, wrecking havoc wherever he went and being despised by many despite his greatness. Likewise, there was a ballplayer whose burning desire to win knew no bounds of civility. Sliding into bases with spikes up and pummeling bothersome spectators was all fair game. Can there be any doubt that Ares is Ty Cobb?
Ares had a counterpart though, a goddess of war who focused on strategy, tactics, and civilized war in contrast to her brother. In baseball, a contemporary of Cobb was in many ways his opposite. A pitcher rather than a hitter; a gentleman rather than a ruffian who would spike his own mother if she was in his way (of course, there was some history there…); and yet, in his own way, with just as much of a will to win. No, he did not spring fully formed from Chadwick’s head, but he might as well have. Athena is Christy Mathewson.
We come to a goddess of the hunt, known for her virginity. The best example of hunting in baseball is the hitter, weapon in hand, staring down the pitcher and attempting to inflict harm upon him. There was one man whose single-minded drive to be a great hitter has been celebrated throughout the years. The virginity angle is a little awkward, but it can be said that you do not hear stories about him chasing women, drinking, or doing anything other than attempting to be the greatest hitter who ever lived. Better yet, he flew a fighter in Korea, engaging in an even more literal hunt. Artemis is Ted Williams.
Artemis’ twin brother is a little harder to relate to his baseball counterpart--a god of song, colonies, medicine, the sun, and more. But there is a player who is celebrated in several songs, who was a great player from a young age, and is forever linked with Ted Williams. Apollo is Joe DiMaggio.
A swift-footed messenger was a little late to the party, but with his cunning, he earned the respect of the other gods. In baseball, an entire race of talented players was cut off from the glory of the major leagues until a fast, fearless, daring player earned his opportunity and earned respect, however belated it was. Hermes is Jackie Robinson.
Then there are those who go to the underworld, whether as punishment or as a neutral afterlife destination. Ruling over them was a god who was an Olympian in his own right, but was born to rule the underworld. In baseball, Hades was sent below as punishment, and so the parallel is note quite exact. The Greeks apparently could not anticipate the pathetic nature of the baseball counterpart. Nonetheless, he exists, and Hades is Pete Rose.
Another god was in charge of fire and forging of weapons. He made a brilliant set of armor for Achilles, arming him for battle. Likewise, a baseball Hephaestus forged weapons for ballplayers. There is no more important tool than a ballplayer’s sword, which he uses to inflict damage on the opposing pitcher, and there is no bat more synonymous with baseball than Louisville Slugger. Hephaestus is John Hillerich.
Unfortunately, the scroll ends there, so we don’t know who the baseball equivalents of Hera, Poseidon, and others are. Nevertheless, maybe next time you sacrifice a chicken to the baseball gods to help the Cubbies, you’ll have a better idea of who it is you are dealing with. Just don’t make Ares mad.
Monday, September 24, 2007
The Baseball Gods
Monday, September 17, 2007
1877 NL
The National League entered 1877 with just six teams after the expulsion of two of its largest markets, New York and Philadelphia. The season would go forward with the six teams, making it the smallest major league ever (a distinction it shares with the 1878 NL and the 1882 American Association). Incidentally, it is along with the 1878 NL the only league in which all of the team nicknames included colors.
There were a number of significant rule and procedure changes enacted for the NL’s second season. For the first time, scheduling was handled by the league office rather than being the responsibility of team management. The schedule was reduced to 60 games from 70 due to the smaller number of teams, but each season series now consisted of twelve games rather than ten.
On the field, home plate was made part of fair territory. But most significant was the elimination of the fair-foul hit (much of the information on this topic comes from the article “The Lost Art of Fair-Foul Hitting” by Robert H. Schafer in The National Pastime #20). Prior to 1877, a ball that hit initially in fair territory was a fair ball, no matter where it bounded after that. In 1877, the new rule defining fair and foul was essentially the modern rule that we are all familiar with.
Before the rule change, a small group of batters had shown proficiency in intentionally batting the ball so that it hit near the plate in fair territory, then bounced off wildly into foul territory. If executed successfully, it almost always resulted in a base hit and often a double. Ross Barnes even hit a fair-foul home run in 1872, the only homer he hit that season.
While the rule changes did bring the treatment of fair and foul closer to modern standards, one antiquated rule that remained in effect was that a foul ball fielded on the first bounce was treated the same as a caught pop-up was. Some observers felt that maintaining this rule while eliminating the fair-foul hit swung the scales too far in favor of the defense, but it would remain in effect off and on for the foreseeable future.
The National League had quite an odd season indeed. The Hartford Dark Blues played all of their home games in Brooklyn, looking to make a better profit; it didn’t work. The Cincinnati Red Stockings, in the midst of another awful season, folded in June; although the franchise was quickly reorganized, they sold Charley Jones, Jimmy Hallinan, and Harry Smith to Chicago during the brief outage. The outcry forced William Hulbert to return Jones to Cincinnati, but the additions did not help the defending champion White Stockings, who tumbled to fifth place, fifteen and a half games off the pace after winning the pennant by six games in 1876.
On August 8, Mike Dorgan of St. Louis supposedly became the first to use a rudimentary catcher’s mask during a game. The innovation was initially mocked, but within a couple short years had become standard equipment, as catchers moved up close to the plate and provided a target for the pitcher. At around that time in the season, Louisville had established themselves at the top of the pack, with a 3 1/2 game lead over Boston on August 13 (the source for the information on the Louisville saga presented here is an article by Daniel E. Ginsburg entitled “The Louisville Scandal” in SABR’s Road Trips). At that point, the Grays embarked on an eastern road trip to Hartford and Boston, going 0-7-1 and yielding the lead to the Red Stockings, who would go on to capture the pennant by seven games over Louisville.
It would later come out that four Grays were involved in a game throwing conspiracy (the exact details have never been fully exposed, in regards to who was the mastermind, if all the players fingered were actually complicit, etc.). They were reserve Al Nichols, shortstop Bill Craver, star left fielder George Hall, and the only man who appeared in the pitcher’s box for Louisville all year, Jim Devlin. There was immediate suspicion and all four were promptly banned for life when the scheme was uncovered.
At the conclusion of the season, the disgraced Louisville franchise exited the league, as did Hartford/Brooklyn, who did not find greener pastures in New York, and St. Louis, leaving uncertainty hanging over the NL as it scrambled to find enough teams to play ball in 1878.
STANDINGS
It would be interesting to see a breakdown on the runs/runs allowed of Louisville before and after the road trip, as Boston comes out far superior to the Grays in component stats, although I suppose that when you have just one pitcher and he’s on the take, it can do a lot of damage to your runs allowed in a short period of time.
So I looked it up. From the eastern trip to the end of the season, the Grays were outscored 86-68 in 20 games. In the eight game eastern trip, they were outscored 38-11 (.110 EW%). If you throw out the last 20 games, they 271 and allowed 202 in 41 games, for a .645 EW%, which while still well behind Boston, shows them to be a much stronger club.
In 1877, the league hit .271/.289/.338 for a .092 SEC, 5.67 runs/game and 24.06 outs/game.
BOSTON
The Red Caps won their fifth pennant in six years after finishing fourth in the inaugural season. A big lift was getting Deacon White back from Chicago, as he is my choice for National League MVP. Tommy Bond was the league’s top pitcher, brought over from Hartford to replace the slightly above average trio employed in 1876, and Jim O’Rourke went from having a fine 1876 to being the NL’s second best position player in 1877.
LOUISVILLE
The Grays were done in by their ace pitcher and by their top position player (Hall). Of the tainted four, only Devlin had been with Louisville in 1876; Hall and Craver were added from the expelled Philly and New York entries. Al Nichols had played for the Mutuals in ’76, and did not play at all in ’77 until added by the Grays in early August after Bill Hague was sidelined by illness.
HARTFORD
The Dark Blues were badly hurt by the loss of both of their highly effective pitchers, as Candy Cummings went to Cincinnati and Bond was the league’s best pitcher in Boston. Rookie Terry Larkin was average as their primary pitcher, and John Cassidy went from playing in twelve games to being the league’s best right fielder. But every other position save first base suffered a noticeable WAR dropoff, and the Dark Blues lost four games relative to the pennant winner.
ST. LOUIS
The Brown Stockings were also crippled by pitcher musical chairs as George Bradley jumped to Chicago, leaving them with a pair of rookie sub-replacement level pitchers in Tricky Nichols and Joe Blong. John Clapp was once again excellent behind the plate, but Joe Battin tumbled from being the circuit’s top third baseman at 153 ARG and +3.1 WAR to a 77 ARG, +.4 WAR campaign. He would not again appear in the major leagues until 1882. The Brown Stockings also lost their top 1876 WAR performer, Lip Pike, to the Red Stockings. On the bright side, Mike Dorgan was the top rookie performers in the game.
CHICAGO
The defending champs took a precipitous dive, much of it traceable to Al Spalding’s decision to give up pitching. George Bradley, signed as Spalding’s replacement and +5.4 WAR with St. Louis in 1876, dropped off to +1, although even a repeat of his good year would have failed to match Spalding’s ’76. Deacon White went back to Boston, and had an MVP season, while Spalding was barely replacement level in taking his spot. Ross Barnes, MVP in ’76, is often cited to have struggled due to the elimination of fair-foul hits. While Barnes was a fair-foul fiend, the more likely culprit in his collapse was a serious, “malaria-like” illness that slowed and then shelved him, threatening his baseball career. To add insult to injury, the White Stockings were the last team in major league history to go the entire season without hitting a home run.
CINCINNATI
If you can ignore the stats and what you know about their 1876 performance, and just look at the names, the Red Stockings look to be a fairly solid team. Candy Cummings, Levi Meyerle, Charlie Gould, Bobby Mathews, Charley Jones, and Lip Pike are all big names. The Red Stockings may have therefore been the first team to field an early 2000s Mets or current Giants-like team that would have been pretty darn good five years earlier. Cummings at 29 was on his last major league legs, as were the 32-year old Meyerle (who would have a brief fling with the questionably major Union Association in 1884), and the 30-year old Gould. Pike at this point was 32 and had one more good season left in him. Jones and Mathews were the only two that had good days in front of them, although Mathews was awful in ’77. Bobby Mitchell, the #3 pitcher, was the first left-handed pitcher in the major leagues.
The Red Stockings were the only team in the NL not to have a winning record at home (12-17); only the champion Red Caps managed a winning record on the road (15-13).
Now the leaders and trailers:
BATTING AVERAGE
1. Deacon White, BOS (.377)
2. John Cassidy, HAR (.378)
3. Cal McVey, CHI (.368)
Trailer: Amos Booth, CIN (.172)
ON BASE AVERAGE
1. Jim O’Rourke, BOS (.407)
2. Deacon White, CHI (.405)
3. Cal McVey, CHI (.387)
Trailer: Will Foley, CIN (.205)
SLUGGING AVERAGE
1. Deacon White, BOS (.545)
2. Charley Jones, CIN/CHI (.471)
3. John Cassidy, HAR (.458)
Trailer: Amos Booth, CIN (.197)
Once again, the Red Stockings swept the trailers in BA/OBA/SLG, but at least poor Redleg Snyder was spared from doing it all himself.
SECONDARY AVERAGE
1. Charley Jones, CIN (.221)
2. Deacon White, BOS (.188)
3. Lew Brown, BOS (.167)
Trailer: Jack Burdock, HAR (.029)
RUNS CREATED
1. Deacon White, BOS (70)
2. Jim O’Rourke, BOS (64)
3. Cal McVey, CHI (60)
4. John Cassidy, HAR (57)
5. George Hall, LOU (56)
ARG
1. Deacon White, BOS (213)
2. Jim O’Rourke, BOS (188)
3. John Cassidy, HAR (187)
4. George Hall, LOU (165)
5. Cal McVey, CHI (160)
Trailer: Will Foley, CIN (45)
WAA
1. Deacon White, BOS (+3.3)
2. Jim O’Rourke, BOS (+2.7)
3. John Cassidy, HAR (+2.4)
4. George Hall, LOU (+2.2)
5. Cal McVey, CHI (+1.8)
Trailer: Will Foley, CIN (-1.7)
WAR
1. Deacon White, BOS (+3.9)
2. Jim O’Rourke, BOS (+3.6)
3. John Cassidy, HAR (+3.3)
4. George Hall, LOU (+3.1)
5. Cal McVey, CHI (+3.0)
Trailer: Will Foley, CIN (-.6)
ARA
1. Tommy Bond, BOS (130)
2. Jim Devlin, LOU (111)
3. Bobby Mitchell, CIN (108)
Trailer: Bobby Mathews, CIN (73)
WAA
1. Tommy Bond, BOS (+3.4)
2. Jim Devlin, LOU (+1.5)
3. Terry Larkin, HAR (+.8)
Trailer: Bobby Mathews, CIN (-1.3)
T WAR
1. Tommy Bond, BOS (+4.1)
2. Jim Devlin, LOU (+3.5)
3. Terry Larkin, HAR (+1.8)
Trailer: Bobby Mathews, CIN (-1.2)
My all-star team:
C: John Clapp, STL
1B: Deacon White, BOS
2B: Joe Gerhardt, LOU
3B: Cap Anson, CHI
SS: John Peters, CHI
LF: Mike Dorgan, STL
CF: Jim O’Rourke, BOS
RF: John Cassidy, HAR
P: Tommy Bond, BOS
MVP: 1B Deacon White, BOS
Rookie Hitter: LF Mike Dorgan, STL
Rookie Pitcher: Terry Larkin, HAR
I gave Clapp the edge over McVey at catcher because Palmer has him at +10 defensively versus McVey’s -10, and the offensive edge is meaningless. Peters was .1 WAR behind Ezra Sutton, but was +20 according to Palmer versus -5, so he gets the nod for the second straight season. Mike Dorgan in left is the choice as George Hall helped tank a pennant, and is thus utterly disqualified in my mind.
Wednesday, September 12, 2007
Welcome to the Club, Josh Newman
Depending on whose count you use, Josh Newman has just become the 47th or 56th Buckeye to play in the major leagues. The 47 figure from the SABR Collegiate Committee and as seen on Baseball-Reference, includes only players who are confirmed to have played for OSU. I was provided some additional names by an anonymous message board poster, which I have included on my website, but these have not been vetted, and probably never will be able to. Unless you can get a full list of all students registered at the school, some old-timers will inevitably come down to "We can't confirm that he went to OSU, but we can't say definitively that he didn't either." Anyway, congratulations to Josh.
Monday, September 10, 2007
1876 NL
I am not a historian by any means, so these yearly recaps will focus mostly on stats and not on interesting tidbits. I will try to write a little bit about the happenings in baseball that season, such as rule changes, special feats like no-hitters, firsts, and the like. These are largely drawn from David Nemec’s Great Encyclopedia of Nineteenth Century Major League Baseball, Total Baseball VI, the ESPN Baseball Encyclopedia, and The Ball Clubs by Donald Dewey and Nicholas Acocella. For more general off-the-field history, the most valuable reference was Harold Seymour’s Baseball: The Early Years, although David Voight’s American Baseball, Vol. 1 is good too. For those of you with knowledge about the nineteenth century game, the narrative will surely be lacking.
The National League was founded at a meeting on February 2 by Chicago White Stockings president William Hulbert. For the previous five seasons, the National Association had been the alliance of professional clubs that would best fit the description of “major league”, although even the NL that supplanted it was far from having the order that we would picture from such a description today. More importantly, even later into the nineteenth century, there were very talented players and excellent teams outside the umbrella of the major leagues. Unlike today, when the thirty major league squads are clearly the best teams in the country, there were likely a large number of minor league teams that were on a fairly equal footing with the NL clubs. And of course it would be many years before the minor leagues were made subservient to the interest of the major league clubs.
Hulbert’s NL put control in the hands of what today we would consider team owners and presidents, unlike the National Association, whose full name was the National Association of Professional Baseball Players. The league also mandated a fifty cent admission charge and no alcohol sales at games. Of course, Hulbert had incentive to form the new league other than extolling the virtues of pleasant, sober crowds and responsible management. The White Stockings would have likely been expelled from the NA for signing players under contract with other teams. Specifically, Hulbert raided four-time defending champ Boston for Al Spalding, Deacon White, Cal McVey, and Ross Barnes. Spalding’s exploits in 1876 were not limited to the diamond; he founded what would become a very lucrative sporting goods retailer bearing his name as well.
The rules of the game had some significant differences from modern baseball. Notable was the square (diamond) home plate, a pitching box situated 45 feet away from the plate, the illegality of overhand pitching, maximums of nine balls and four strikes, batters calling for a high or low pitch, no free trip to first base for getting plunked by a pitch, the catcher standing rather than squatting behind the plate, and so-called fair-foul hits. If a ball hit in fair territory before passing the corner bases, and then went foul, it was fair. The aforementioned Barnes was a well-known practitioner of this art.
On April 22, the first NL game was played, and the Boston Red Caps, shorn of their stars, beat the Philadelphia Athletics 6-5 in Philadelphia. Joe Borden got the win for Boston, a second notable feat for him--in 1875, he tossed the first NA no-hitter. It was George Bradley of the St. Louis Brown Stockings who would get the NL’s first no-hitter, beating the Hartford Dark Blues 2-0 on July 15. On September 9, Candy Cummings of Hartford beat the Cincinnati Red Stockings twice in the NL’s first doubleheader.
The pennant race was not particularly notable, as Hulbert’s White Stockings went 52-14 to finish six games ahead of both the Brown Stockings and the Dark Blues. More significant may have been Philadelphia and New York’s refusal to make their respective late-season western road trips, given their being buried in the second division. Such practices were common in NA days, but Hulbert was able to engineer their expulsion from the league, demonstrating that the NL would have a much stronger central office than the NA, and that the rules were not just suggestions.
STANDINGS
As you can see, the two W% estimators are often at odds with each other; I would again caution against reading too much into PW% here. They all agree that Chicago was the best and Cincinnati was hapless. One aside here is that a team like Cincinnati, with a .138 W% and .136 EW%, cannot possibly be evaluated successfully by the Win Shares methodology. While there are no such extreme modern teams, Win Shares were published for the 19th century. I would be very skeptical about using them at all, but I would heed them no attention for a team like the Red Stockings.
Here are each team’s primary players at each position along with some important sabermetric stats. If a team used multiple pitchers for serious innings, I have listed both of them. WAA is against an average hitter, regardless of position; WAR is against a replacement level hitter at the position of the player in question. T WAR for pitchers is their offensive WAR plus their pitching WAA. I may at times make reference to so and so being the MVP, or the all-star shortstop, etc. These are based on my choices, not official awards, as there were none.
In 1876, the league hit .265/.277/.321, for a .072 Secondary Average, with 24.22 outs/game and 5.90 runs/game.
CHICAGO
The White Stockings were clearly baseball’s best team, and the four players they raided from Boston were a huge part of that. After finishing 35 games behind Boston in 1875, they finished 15 games ahead of them in 1876. White was the #2 catcher by WAR, McVey the #2 first baseman, Barnes the best player period, and Spalding the game’s top pitcher. The other Chicago regulars were holdovers from their 1875 team, with the exception of Bob Addy, picked up from a defunct Philadelphia outfit. In the over 130 seasons that have followed, only one NL entry has topped their winning percentage. With the flag, representatives of Chicago had won the first NL championship; twenty-five years later, another Chicago team known as the White Stockings would take the first AL pennant.
HARTFORD
Of the eight NL teams, four made extensive use of multiple pitchers. But the Dark Blues were the only team with the luxury of two good pitchers in Bond and Cummings. RF Dick Higham will reappear in the narrative in later years, although not for positive reasons. Bob Ferguson is the holder of the best encyclopedia nickname ever, “Death to Flying Things”. Along with Chicago, Hartford was the only team to boast an above average hitter at each lineup position.
ST. LOUIS
BOSTON
How would the Red Stockings have done with their four stars still in tow? Well, Deacon White was 2.5 WAR better than Lew Brown, Cal McVey 1.1 better than Tim Murnane, Barnes a whopping 4.6 better than John Morrill, and Spalding 5.4 better than his three-headed replacement. That’s 13.6 in total, and Chicago’s margin over Boston was 15 games, so it’s pretty clear that Hulbert’s raid was directly responsible for a pennant (not that this is a startling new insight).
According to The Ball Clubs, Joe Borden had been signed to a three year contract. The president of the Red Caps, seeking to get Borden to beg for a buyout of his contract, put him to work as a groundskeeper at the club’s field. His plan backfired when Borden went about his new job, leaving them with a very expensive groundskeeper, and forcing the Red Caps to offer Borden a generous buyout, which he accepted. Borden also earned the tongue-in-cheek sobriquet “Josephus the Phenomenal” during his Boston career.
LOUISVILLE
I don’t have anything smart to say about them, but I think I’ve worked in the nickname for each of the teams somewhere except for the Grays. So there you have it. The format of [City Name] [Nickname] was far from standard at this time, but most teams did have at least an informal tag that was used. I have usually gone with the name that Nemec used, but one should always keep the fact that this is somewhat of a modern concept being forced upon a previous era. Reserve outfielder George Bechtel was expelled from the league by the Grays for attempting to throw a game and trying to bribe teammates to join him in that effort.
NEW YORK
With just one above average hitter and one replacement-level pitcher who recorded all but one of the team’s decisions, the Mutuals can scarcely be blamed for failing to play out the string.
PHILADELPHIA
The other quitters are a little harder to figure, as both their expected and predicted W%s point to a bad team that should have ended up much better than 14-45. I imagine that their defensive innings were fun to watch, as #2 pitcher George Zettlein allowed the other team to put the ball in play at a very high rate, even for the time and place. David Voigt, in American Baseball Vol. 1, reports that “few pitchers were more accident prone than George Zettlein. He was hit hard and often by batted balls; once a reporter saw a line ball hit him with such force that the ball rebounded sixty feet. Somewhat stunned, Zettlein ‘shook his head, took a drink, and again went to work as if nothing had happened.’” And than down at third base was Levi Meyerle, a man who was born a century too early to benefit from the DH. The Athletics had ten catchers appear in games, which has to be some kind of record, although I am too lazy to check. Malone, Fisler, and Meyerle were holdovers from the first pro pennant winners, the NA’s 1871 Athletics.
CINCINNATI
The only above average hitter on this sorry bunch was Charley Jones, a colorful character with an unknown demise. Redleg Snyder at least had an appropriate sobriquet for a member of this organization, although this is not the same Reds franchises that survives to this day. In fact, only the Boston Red Caps (later the Braves, and of course now in Atlanta) and the Chicago White Stockings (later the Cubs), still play in the major leagues today.
Now let’s look at the league leaders and trailers in some useful categories:
BATTING AVERAGE
1. Ross Barnes, CHI (.429)
2. George Hall, PHI (.366)
3. Cap Anson, CHI (.356)
Trailer: Redleg Snyder, CIN (.151)
ON BASE AVERAGE
1. Ross Barnes, CHI (.451)
2. George Hall, PHI (.476)
3. Cap Anson, CHI (.380)
Trailer: Redleg Snyder, CIN (.155)
SLUGGING AVERAGE
1. Ross Barnes, CHI (.590)
2. George Hall, PHI (.476)
3. Lip Pike, STL (.472)
Trailer: Redleg Snyder, CIN (.176)
Redleg Snyder pulls off a dubious triple crown, although in a game with few walks and little power, it’s not surprising that the guy with the lowest BA would also have the lowest OBA and SLG.
SECONDARY AVERAGE
1. Ross Barnes, CHI (.224)
2. George Hall, PHI (.209)
3. Lip Pike, STL (.177)
Trailer: Mike McGeary, STL (.018)
RUNS CREATED
1. Ross Barnes, CHI (98)
2. Cap Anson, CHI (71)
3. George Hall, PHI (69)
4. Jim O’Rourke, BOS (66)
5. John Peters, CHI (66)
ARG
1. Ross Barnes, CHI (227)
2. Lip Pike, STL (192)
3. Dick Higham, HAR (165)
4. Jim Devlin, LOU (162)
5. Joe Battin, STL (153)
Trailer: Redleg Snyder, CIN (32)
WAA
1. Ross Barnes, CHI (+4.1)
2. Lip Pike, STL (+3.1)
3. Dick Higham, HAR (+2.4)
4. Jim Devlin, LOU (+2.2)
5. Joe Battin, STL (+1.9)
Trailer: Redleg Snyder, CIN (-2.1)
WAR
1. Ross Barnes, CHI (+5.3)
2. Lip Pike, STL (+4.1)
3. Jim Devlin, LOU (+4.1)
4. Dick Higham, HAR (+3.5)
5. John Clapp, STL (+3.3)
Trailer: Redleg Snyder, CIN (-1.2)
ARA
1. Al Spalding, CHI (173)
2. Tommy Bond, HAR (138)
3. George Bradley, STL (133)
Trailer: Dory Dean, CIN (68)
WAA
1. Al Spalding, CHI (+4.8)
2. George Bradley, STL (+3.2)
3. Tommy Bond, HAR (+2.5)
Trailer: Dory Dean, CIN (-2.7)
T WAR
1. Al Spalding, CHI (+7.3)
2. George Bradley, STL (+5.4)
3. Jim Devlin, LOU (+4.2)
Trailer: Dory Dean, CIN (-2.1)
Hmm…Chicago had the highest individual player and pitcher WARs and they won the pennant. Cincinnati had the lowest individual player and pitcher WARs and they finished in the cellar. Funny how that works.
At this point, let me pick an all-star team, largely based on WAR, but maybe throwing a bit of the players’ defensive reputations if appropriate. The all-star pitcher would be the “Cy Young” winner, although Denton was nine so I don’t think they would have called it that.
C: Deacon White, CHI
1B: John Clapp, STL
2B: Ross Barnes, CHI
3B: Cap Anson, CHI/Joe Battin, STL
SS: John Peters, CHI
LF: George Hall, PHI
CF: Lip Pike, STL
RF: Dick Higham, HAR
P: Al Spalding, CHI
MVP: 2B Ross Barnes, CHI
Rookie Hitter: CF Charley Jones, CIN
Rookie Pitcher: Foghorn Bradley, BOS
The Anson/Battin split is because their numbers are nearly identical. Anson obviously has the name that is recognizable today, but in 1876 I have Anson at 152 ARG (52% above his contextual average), Battin at 153. So Anson winds up with 1.82 WAA and 3.07 WAR, and Batting has 1.85 and 3.09. Those differences are too small to be meaningful, so I checked Pete Palmer’s Fielding Runs figures on the two of them. I don’t have a lot of confidence in FR (or most other fielding metrics that don’t have the benefit of PBP data), mind you, even more so in the 1800s, but I’m sure they do a decent job of picking out the Meyerles. Battin is at +14 FR, Anson +13. They are fractions of a run apart on both offense and defense; I think that’s a tie.
The ESPN Encyclopedia gave their ex post facto Cy to George Bradley. The ex post facto awards are not an attempt on their part to say who should have won, but rather who likely would have won had the award existed. Anyway, Bradley certainly turned in a more dominant defense independent performance, as he had a 103/38 K/W in 573 IP, while Spalding was just 39/26 in 529 IP. It may well be that Bradley deserves more credit for his performance, but I have not attempted to go too deep with pitcher evaluation here.
Two of my all-stars will eventually be banned from baseball for life, raising legitimate questions about the quality of their effort in this year, but we’ll pretend not to know that. And the star of stars, MVP Ross Barnes, is about to have his career crash and burn, but we don’t know that either.
Monday, September 03, 2007
Early NL Series: Pitchers and Teams
We have looked at batter evaluation; now we need to tackle pitchers. Or we could run away whimpering and go hide in the closet. This at times seems like a very appealing option when you have to deal with the problems of evaluating early (in baseball terms) nineteenth century pitchers.
First, throw win-loss records out the window. Many pitchers are throwing most of their team’s innings, making them almost solely responsible for the team’s W% (insofar as the pitching staff can control it), which makes it impossible to compare win-loss records to those of the team. Of course, even if that was feasible, it would not be preferable. I am an advocate of looking at win-loss records intelligently, but only on a multi-season level (and still with many caveats), and we want to be able to evaluate single seasons.
What about Run Average? Well, there are so many errors that even a die-hard RA supporter like me starts to get queasy about using it. Not to mention that fielding was a much bigger factor in the game then it is today. In modern times, when we evaluate a pitcher solely based on RA, we may overstate his value a bit--after all, some of the credit is due to the fielders. But in this time, we will be wildly overstating his value.
As for ERA, I oppose the practice of “reconstructing innings” when there is one error per game, as I think it muddies things up more then it clarifies them, by wiping out real events given up by the pitcher just because they were preceded by an error. How much worse then would it be when errors infest the game? So many innings were reconstructed to produce earned runs that we could be losing tons of information.
How about an Estimated RA, most often exemplified by Bill James’ Component ERA? It will allow us to use expected errors in place of actual errors, but it still will not alleviate the problem of over-crediting the pitcher.
So then we come to DIPS. But we quickly realize that the three true outcomes are a misnomer in this time; most home runs are actually in play. We can’t set them aside, and even if we did, they would have very little effect on the evaluation as there were only 533 home runs hit over the eight seasons in question, just 1.2% of all hits. And strikeouts and walks make up just 10% of all plate appearances, compared to, just picking a year out of the air, 17.4% in the 1939 NL. That doesn’t leave us with a large base to evaluate on.
And who’s to say that hits/ball in play were as bunched together as they are now? Given the greater importance of fielding, they may not have been. On the other hand, pitchers are throwing underhand from 45 feet, needing a large number of strikes to retire the batter, and without the wide variety of pitches we see today. So maybe we should expect all hurlers to be about equal in that regard.
If we try to test year-to-year correlations, though, we will have a lot of issues. Most obvious is the small sample size. Of course selective sampling is ever-present. We can’t really compare a pitcher to his teammates, since he often makes up such a large share of the team total. Even if we were able to and found a low year-to-year correlation on H/BIP, that would not necessarily render it meaningless or make a DIPS-like approach appropriate for evaluating value retrospectively.
So what to do? I have decided to go an extremely lazy route, and to evaluate pitchers based on Wins Above Average, based on runs allowed, except I have multiplied that result by the league percentage of earned runs for the given season. This is a terrible hodge-podge of an approach, as I don’t think that earned runs are particularly useful. But you have to account in some way for the fact that fielding was a much larger share of defense then it is today, and my guess would not be any more legitimate then the percentage of earned runs.
I have also chosen average as a baseline, rather then replacement, because I have decided to evaluate the pitchers first as hitters, versus a replacement level hitter, and then add on their pitching value. This is what I would do today for a position player, at least for simplicity’s (there are many knocks against offensive positional adjustments and I don’t disagree) sake; first find the batter’s offensive value versus an average player at his position, and then add in the runs he saved above an average fielder at his position.
For an example of how this works, let’s take a look at 1876 Hartford’s #2 pitcher, Candy Cummings of supposed curveball invention fame (the Dark Blues played 69 games; Tommy Bond started and completed 45 of them, Cummings 25). The Dark Blues context is 10 total RPG, so we would expect Candy to allow 5 runs per 9 innings. In fact, he allowed 97 runs in 216 innings for a 4.04 RA. This means that his Adjusted RA, which I will use pretty extensively instead of the actual figure, was 100*5/4.04 = 124.
In 1876, 39.6% of the league runs were earned. Cummings’ RAA is therefore (5-4.04)*.396*216/9 = +9.12. Converting this to WAA is as simple as dividing by 10, and so he is +.91 wins. If you want formulas:
ARA = 200*RPG/RA
RAA = (RPG/2 - RA)*LgER%*IP/9
WAA = RAA/RPG
Cummings was a “replacement-level hitter” for a pitcher (the phrase “replacement-level hitter” is silly, since what we really mean by a replacement level player is one who is replacement level in his total value; however, it is a necessary evil when one has taken the flawed offensive positional adjustment path), with +.01 WAR, so his total value is +.92 WAR.
As for teams, things are much more straightforward. Expected Winning Percentage based on runs and runs allowed can be calculated using Pythagenpat. I have also figured Predicted W%, based on Base Runs and Base Runs Allowed; this is a little more shaky since the Base Runs formula isn’t particularly accurate, and many of the defensive components had to be estimated. To find PW%, just use BsR and BsRA as you would R and RA. I have included those figures here, but I wouldn’t put much stock into it for this era.
Monday, August 27, 2007
Early NL Series: Batter Evaluation
We have a Runs Created formula, the backbone of just about any offensive evaluation system. But what about the other details, like park adjustments, performance by position, baseline, and conversion to wins?
I’ll deal with these one-by-one. Park adjustments in this time would be a real pain to figure. Parks change rapidly, they may have the same name but be a new structure from year-to-year, teams are jumping in and out of the league (which radically alters the “road” context for any given team from year-to-year), sample size is reduced because of shorter seasons, etc. I suppose that I could look at all of this and say it’s not worth doing the work, and just use Total Baseball’s PFs. But that is not the option I have chosen, because Total Baseball’s park factors are subject to the same problems that the ones I would calculate would be.
So instead I have decided to simply use each team’s actual RPG as the standard to which a hitter is compared. This is problematic in some sense, if your goal is building a performance or ability metric--a player on one team may be valued more highly because his team has a good pitching staff, while another may be hurt if the opposite holds. On the other hand, from a value perspective, as Bill James argued way back in 1985, the result of other games don’t define the player in question’s value--his contributions to wins and losses comes only in the context of the games that his team actually plays.
Additionally, there is a huge gorilla in the room in the whole discussion of early major league baseball that I am ignoring, and that is fielding. Fielding is a pain to evaluate in any time, and it is not my specialty in any case, so I am not going to even begin to approach an evaluation of the second baseman in the 1879 NL. So of course my ratings will only cover offense and while you should take offense-only ratings with some skepticism in today’s game, it is even more so in a game where the ball is put in play about 90% of the time and there are 5 errors/game, etc. But by using the team’s actual RPG figure, we do capture some secondary effects of, if nothing else, the whole team’s fielding skill (less runs allowed means a lower baseline RG for players, as well as less runs per win). In the last paragraph I slipped in the argument about a “good pitching staff” potentially inflating a batter’s value. But in this game it is even harder to separate pitching from fielding than it is today, and that “good pitching staff” is more likely “good pitching and fielding”, to which our player has contributed at least a tiny bit. This is not in an attempt to say that the RPG approach would be better than using stable PFs if they existed, but just a small point in its favor that would be less obvious in 2006.
I am still going to apply a positional adjustment to each player’s expected runs created, based on his primary position. Looking at all hitters 1876-1883, the league had a RG of 5.36. Here are the RG and Adjusted RG(just RG divided by the prevailing RPG, in this case 5.36) for each position:
This actually matches the modern defensive spectrum if you throw out the quirky rightfield happenings. One potential explanatory factor I stumbled upon after writing this came from an article by Jim Foglio entitled “Old Hoss” (about Radbourn) in the SABR National Pastime #25:
“When Radbourn did not start on the mound, he was inserted into right field, like many of his contemporary hurlers. In 19th-century ball, non-injury substitutions were prohibited. The potential relief, or ‘exchange’ pitcher, was almost always placed in right, hence the bad knock that right fielders have received all the way down to the little league game…[the practice] surely had its origins in [pitchers’] arm strength, given the distance of the throws when compared to center and left.”
However, I’m not sure this explains the right field problem satisfactorily, because I did not look at offensive production when actually playing a given position; it was the composite performance of all players who were primarily of a given position. Change pitchers out in RF, if they got in more games as pitchers than as right fielders, would still be considered pitchers for the purpose of the classifications used to generate the data above. Of the men classified by Nemec as the primary RF, in the 1876-1883 NL there are 60 players. Of those 60, 23 (38%) pitched at some point during that season, but only 8 (13%) pitched more than 30 IP.
Another possibility is that if left-handed hitters are more scarce, there are fewer balls hit to right field. This might actually be the most satisfying explanation; I did not actually check the breakdown on left-handed and right-handed batters, but it figures that there were less lefties than in the modern game. I do know for a fact that there were very few left-handed pitchers at this time.
Excepting right field, the degree of the positional adjustments is not the same as it is today, but the order is essentially the same. Notably, Bill James has found that at some point around 1930, second and third base jumped to their current positions on the spectrum. But here, third baseman are still creating more runs than second baseman. Of course, the difference is not that large, but it is possible that a later change in the game caused the jump, and then in 1930 it just jumped back to where it had been at the dawn of the majors. On the other hand, it could just be an insignificant margin or a result of a poor RC formula, or what have you. Perhaps the increase in bunting, particularly in the 1890s, made third base a premium defensive position, and the birth of home run ball in the 20s and 30s and the corresponding de-emphasis of the bunt changed that. And of course it is always possible, as is likely the case for RF, that the offensive positional adjustment does not truly reflect the dynamics in play, and is a poor substitute for a comprehensive defensive evaluation, or at the very least defensive positional adjustment. What I have done is treat all outfielders equally, using the overall outfield PADJ for the WAR estimates below.
Then we have the issue of baseline. I am reluctant to even approach it here, since it always opens up a big can of worms, and most of the ways you can try to set a “replacement level” baseline are subject to selective sampling concerns. But I went ahead and fooled around with some stuff anyway. In modern times, some people like to use the winning percentages of the league’s worst teams as an idea of where the replacement level is. In this period, the average last-place team played .243 ball, while the second-worst was .363 and the average of the two .303. So this gives us some idea of where we might be placing it.
I also took the primary starting player at each position for each team, and looked at the difference in RG between the total, the starters, and the non-starters. I also tossed out all pitchers. All players created 5.37 RG, while the non-pitchers were at 5.55. Starting non-pitchers put up 5.78, while the non-starters were 3.68. Comparing the non-starters to all non-pitchers, we get an OW% of .305 (while I hate OW%, it is common standard for this type of discussion). This is very close to the .303 W% of the two worst teams in the league (yes, this could be due to coincidence, and yes, I acknowledge selective sampling problems in the study). And so I decided to set the offensive replacement level at .300.
In our day I go along with the crowd and use .350--despite the fact that I think various factors (chaining and selective sampling chief among them) make that baseline too low. Here I’ve gone even lower, based on the same kind of faulty study. Why? Well, this is subjective as all get out, but in this time, in which there is no minor league system to speak of, when you have to send a telegraph to communicate between Boston and Chicago, when some teams come from relatively small towns by today major league standards, and when the weight of managerial decision-making was almost certainly heavy on fielding than it is today, I think that a lower baseline makes sense. If Cleveland loses one of its players, it doesn’t have a farm club in Buffalo it can call and get a replacement from. They may very well just find a good local semi-pro/amateur/minor league player and thrust him into the lineup.
On the other hand (I’m using that phrase a lot in the murky areas we are treading in), the National League, while clearly the nation’s top league, is not considered the only sun in the baseball solar system as it is now. Many historians believe that the gap between a “minor” league and the National League was much, much smaller than it would be even in the time of Jack Dunn’s clubs in Baltimore and a strong and fairly autonomous Pacific Coast League. So perhaps a .300 NL player is really just not a very good ballplayer at all, and a team stuck with one could easily find a better one, if only there were willing to pay him or find him somewhere.
It would take a better mathematical and historical mind than I have to sort this all out. I am going to use .300, or 65% of the league RG, as the replacement level for this time.
So now the only issue remaining in the evaluation of batters is how we convert their runs to wins.
In 1876, Boston’s RPG(runs scored and allowed per game) was 13.16. Pythagenpat tells us that the corresponding RPW is 12.47. In the same league, Louisville’s RPG was 9.04, for 9.55 RPW. This is a sizeable difference, one that we must account for. A Boston player's runs, even compared to a baseline, are not as valuable as those of a Louisville player.
However, my method to account for this will not be as involved as the Pythagenpat. I will simply set RPW = RPG. This has been proposed, at least for simple situations, by David Smyth in the past. And as Ralph Caola has shown, this is the implicit result of a Pythagorean with exponent 2. That is not to say that it is right, but when we have all of the imprecision floating around already, I prefer to keep it simple. Also, it doesn’t make that much of a difference. The biggest difference between the two figures is .37 WAR, and there are only twelve seasons for which it makes a difference of .20. For most, it is almost as negligible of a difference as an extra run created.
Let me at this point walk you, from start to finish, through a batter calculation, so that everything is explained. Let’s look at Everett Mills, 1876 Hartford, who ends up as the #3 first baseman in the league that season.
Mills basic stats are 254 AB, 66 H, 8 D, 1 T, 0 HR, 1 W, and 3 K. First, we estimate the number of times he reached on an error. In 1876, the average was .1531 estimated ROE per AB-H-K (as given in the last installment). So we give Mills credit for .1531*(254-66-3) = 28.3 errors. This allows us to figure his outs as AB-H-E, or 254-66-28.3 = 159.7.
Then we plug his stats into the Runs Created formula as given in the last installment:
RC = .588S + .886D + 1.178T + 1.417HR + .422W + .606E - .147O
= .588(66-8-1-0) + .886(8) + 1.178(1) + 1.417(0) + .422(1) + .606(28.3) - .147(159.7) = 35.9 RC.
Next we calculate Runs Created per Game, which I abbreviate as RG. The formula is RG = RC*OG/O, where OG is the league average of estimated outs per game, and O is the estimated outs for Mills(159.7 as seen above). In 1876, there are 24.22 OG, so Mills’ RG is 35.9*24.22/159.7 = 5.44.
Now we are going to calculate his Runs Above Replacement, specifically versus a replacement first baseman, who we assume creates runs at 121% of the prevailing contextual average. What is that contextual average? Well, it is half the RPG of Mills’ team. Hartford scored 429 and allowed 261 runs in 69 games, so their RPG is (429+261)/69 = 10. Half of that is 5.
So we expect an average player in Mills’ situation to create 5 runs. But an average first baseman will create 21% more than that, or 1.21*5 = 6.05. So we rate Mills as a below average hitter for his position. But remember that we assume a replacement player will hit at 65% of that, which is .65*6.05 = 3.93. So Mills will definitely have value above replacement; in fact 5.44-3.93 = 1.51 runs per game above replacement.
And Mills made 159.7 outs, which is equivalent to 159.7/24.22 = 6.59 games, so Mills is 1.51*6.59 = 9.95 runs better than a replacement level first baseman. If you want that all as a drawn out formula:
RAR = (RG - (RPG/2)*PADJ*.65)*O/OG
= (5.44 - (10/2)*1.21*.65)*159.7/24.22 = 9.95
Now we just need to convert to Wins Above Replacement. WAR = RAR/RPG. So 9.95/10 = 1.00 WAR.
So our estimate is that Everett Mills was one win better than a replacement level first baseman in 1876. At times, when I get into looking at the player results, I may discuss WAR without the positional adjustment or Wins Above Average, with or without the positional adjustment. If there’s no PADJ, just leave it out of the formula or use 1. If it’s WAA, leave out the .65.
Also, I may refer to Adjusted RG, which will just be 200*RG/RPG.
One little side note about Mr. Mills. You may have noticed that I said he is the third-best first baseman in 1876 in WAR, but is in fact rated as below average. Granted, there are only eight teams in the league, but we would still expect the third-best first baseman to be above average. In fact, only two of the 1876 first baseman have positive WAA, and the total across the eight starters plus two backups who happen to be classified as first baseman (with a combined total of just 29 PA) is -5.1. Apparently, these guys hit nowhere near 121% of the league average. In fact, they had a 6.03 RG v. 5.90 for the league, only a 102 ARG. Whether this is just an aberration or if first baseman as a group had not developed their hitting prowess yet, I have not done enough checking to hazard a guess. It is clear though, by the 1880s, with the ABC trio (Anson, Brouthers, and Connor) that first base was where the big boppers played.
During one of my earlier abortive attempts to analyze National Association stats, I found that in 1871 at least, first base was one of the least productive positions on the diamond. Did first base shift, position on the spectrum, or are we dealing with random fluctuations or the inherent flaws in offensive positional adjustment? Historians are probably best equipped to address this conundrum.
Monday, August 20, 2007
The Audacity of OPS
I have intended for some time to write a post or a series of posts discussing all of the various means of combining OBA and SLG into a more complete stat that are floating around out there. This is not that piece; I have no intentions of discussing all the various OBA/SLG formulas or discussing the technical implications of the ones I do discuss. This is more of a rant, to get this off of my chest.
The title is a little misleading; OPS isn’t really audacious, although perhaps some of its supporters do fall into “recklessness” which is one definition my dictionary gives. I was just too impressed with myself for coming up with the clever phrase (do you have to be a political junkie to get it? More likely, it’s bad and even if you do get it, it’s more likely to elicit a groan then a chuckle).
I should also make it clear before I start: OPS is not a bad statistic. It certainly beats the pants off of looking at the triple crown stats, and there is absolutely nothing wrong with using OPS for quick comparisons or for studies involving large groups of players, or any such thing. But I do believe it is important to keep in mind what OPS is and is not--it is a decent, quick way to evaluate a hitter. But it is not a stat denoted in any sort of useful unit or estimated unit; it is not a stat that was constructed based on a theory about how runs are scored; and it is not a stat that you should go out of your way to use if you have other alternatives available.
First, let’s talk about the components of OPS themselves. On Base Average is a fundamental baseball measurement, because it is essentially the rate of avoiding batting outs. You don’t need me to explain to you why avoiding outs is so important, and that is the point. OBA is a statistic that we would want to invent if we did not already happen to have it sitting around.
Slugging Average is not a fundamental baseball measurement. SLG may be fairly intuitive, and it certainly is venerable, but it is not something that obviously is an important measurement to have on its own. After all, slugging average doesn’t really measure power, because it includes singles. So then what does it measure? It is bases gained on hits by the batter per at bat. But what is the greater significance of bases gained by the batter per at bat?
It really has none. Certainly it is good for batters to gain bases on hits; but that, in and of itself, is not a meaningful measurement. You can even look at the game in such a way that the goal is to gain bases--but in that case, the goal is not for the batter to gain bases, it is for the team to gain bases. And a team doesn’t gain one base for each single on average, nor four bases for each homer, nor do the ratios between one base for a single, two for a double, etc. hold when talking about the bases gained by the team.
The point is not that Slugging Average is meaningless or stupid; the point is that it just is. It is one way of attempting to quantify the value of hits other then counting them all equally as batting average does. It is a crude way of doing so, but it does have a fairly strong correlation with runs and it is a nice thing to know.
But if we didn’t already have SLG, would somebody have to invent it? I think not. There are ways to use the same inputs that would be more useful and would better reflect the run creation value of those inputs.
My message in all of that is that it is not as if OBA and SLG are both (again, my claim is that OBA is, SLG is not) obvious, fundamental things that you would want to know about a team or player’s hitting. They just happen to be two statistics that are the most telling of the widely-available stats.
And it just so happens that when you add them together, you get a measure that is very highly correlated with runs scored, easy to explain computationally, and widely accessible due to the proliferation of OBA and SLG data.
But if you did not have OBA and SLG available to you, would you think of going about creating them so that you could add them together into some uberstat? I would certainly hope not. And how simple is OPS, really? It is simple to compute, in a way, and it is simple to explain, but is it simple to explain why those two things should be added together other then that “it works”? If somebody asks you, “How do you know that just adding them together weights them properly?”, how do you respond?
I said above that OPS is simple to compute, in a way. What I meant by this is that OPS is simple to compute, if you alrseady have OBA and SLG computed for you. Then it is just a simple addition. What if somebody has not already computed them for you? Well now you have (H+W)/(AB+W) + TB/AB, which is not nearly that simple, and not a whole lot simpler then (TB+.8H+W-.3AB)*.32/(AB-H), which gives you a much better rate stat (runs created per out). I guess it can be said that it avoids multiplication--if somebody has already figured total bases for you. If you have to do that yourself, now you have (H+W)/(AB+W) + (H+D+2T+3HR)/AB, which is not that much more simple then (1.8H+D+2T+3HR+W-.3AB)*.32/(AB-H). I guess it can be said that it only includes whole coefficients, if you want to argue on its behalf.
What if we think about what OPS looks like if you write it with a common denominator? Now we have:
OPS = ((H+W)*AB + TB*(AB+W))/(AB*(AB+W))
Not so simple anymore, and can anyone possibly explain the logic behind multiplying those things together like that, other then that “it works”?
Then there is the matter of OPS+. Some people are really shocked to learn that OPS+ is calculated as OBA/LgOBA + SLG/LgSLG - 1. “This is not a true relative OPS!”, they exclaim. “It doesn’t really mean that a 100 OPS+ batter was 20% better in OPS then an 80 OPS+ batter!” While these statements are true, and there is a legitimate complaint to be lodged about the naming of the statistic, the horror at the sacred construction of OPS being violated is somewhat audacious.
When somebody tells you that a stat is adjusted, or has a “+” suffix on the end of it, you expect it to be the ratio of the player’s stat to the league average, perhaps with a park adjustment thrown in. You don’t expect it to be a similar but different statistic. So the measurement that is labeled OPS+ does mislead. Give Pete Palmer a slap on the wrist for this, and move on.
Then when you do move on, give Pete Palmer a pat on the back. Why? Simply for the fact that the measure he has given you is more telling then the measure that you were expecting. I’m not going to get into the math here, since I plan on covering that in my later series, and I will ask you to take this on faith (at least with regards to what is presented here; this is not new information and it has been shown by other sabermetricians in other places). Let’s call OPS/LgOPS “SOPS+” for “straight OPS” plus. People think that OPS+ is SOPS+, and it is not.
The real effect of OPS+, other then adjusting for the league average, is to give more weight to the OBA portion of OPS. Not sufficiently enough weight, but around 1.2 times as much as it is given under OPS or SOPS+. The other thing it does, in addition to correlating better with run scoring, is to express itself in a meaningful estimated baseball unit. OPS+ can be viewed as an approximation, an estimate, of runs per out relative to the league average, which is what you really want to know (or at least is a lot closer to what you really want to know then the ratio of OPS to lgOPS is). Since OPS is unitless, SOPS+ is unitless as well. You can of course use SOPS+ to approximate relative runs/out as well. However, in order to do it, you have to take two times SOPS+, minus one.
So when people complain that OPS+ distorts the ratio between player’s OPS, they are right. But this distortion is a good thing, since it puts it in terms of a meaningful standard instead of a ratio of a contrived, not theoretically-based statistic (OPS). SOPS+ wouldn’t tell these folks what they think it would. A 120 SOPS+ hitter would be 20% above the league average in OPS. That does not in any way, shape, or form mean anything other then that. It does not mean that they created 20% more runs per plate appearance then an average player. It does not mean that they created 20% more runs per out then an average player. It does not mean that they were 20% better then an average player. It does not mean that they are 20% more talented then an average player.
To me, it is a parody of sabermetrics when people complain about OPS+ not being SOPS+, for any reason other then the confusion caused by its name. We sabermetricians have used OPS, and now we will complain about something that is no longer pure OPS, even though it is a more meaningful statistic with clearer units that correlates better with wining baseball games. What is inherently superior about adding OBA to SLG and then comparing to the league average versus comparing OBA to the average, SLG to the average, and adding the results? Considering that OPS doesn’t have any units to begin with and doesn’t correlate better with runs scored, nothing that I can see.
I apologize for the rambling nature of this, but I warned you when I started that it was a rant. OPS is a fine, quick way to measure a hitter. That does not mean that its units are meaningful, that does not mean that it is has meaningful units when it is divided by the league average, or that it is a statistic that has any inherent logic behind it other then adding together two things because it works, or that another metric that combines OBA and SLG in a different way is necessarily inferior or incorrect. As long as you keep those things in mind, there’s not really anything audacious about OPS.
Monday, August 13, 2007
Early NL Series: Intro and Run Estimation
My major interest in baseball research is theoretical sabermetrics. The “theoretical” label sounds a bit arrogant, but what I mean is that I am interested particularly in questions of what would happen at extremes that do not occur with the usual seasonal Major League data that many people analyze (for instance, RC works fine for normal teams, and so does 10 runs = 1 win as a rule of thumb. You don’t really need BsR or Pythagenpat for those types of situations--they can help sharpen your analysis, but you won’t go too far off track without them.) Thus my interest in run and win estimation at the extremes, as well as evaluation of extreme batters (yes, I still have about five installments in the Rate Stat series to write, and yes, I will get around to it, but when, I don’t know). Secondary to that is using sabermetrics to increase my understanding of the baseball world around me (example, how valuable is Chipper Jones? What are the odds that the Tigers win the World Series? Who got the better of the Brewers/Rangers trade?). I don't do this a whole lot here because there are dozens and dozens of people who do that kind of stuff, and I wouldn't be able to add any added insight. But a close third is using sabermetrics to evaluate the players and teams of the past. Particularly, I am interested in applying sabermetric analysis to the earliest days of what we now call major league baseball.
A few years ago, and again recently, I turned my attention to the National Association, the first loose major league of openly professional players that operated from 1871-1875. However, this league, as anyone who has attempted to statistically analyze it will know, was a mess. Teams played 40 games in a season; some dropped out after 10, some were horrifically bad, Boston dominated the league, etc. All of these factors make it difficult to develop the kind of sabermetric tools (run estimators, win estimators, baselines) that we use in present day analysis. So I finally threw my hands up and gave up (Dan Rosenheck came up with a BsR formula that worked better for the NA then anything I did, but there are limitations of the data that are hard to overcome). For now, it is probably best to eyeball the stats of NA players and teams and use common sense, as opposed to attempting to apply rigorous analytical structures to them.
Anyway, when things start to settle down, you have the National League, founded in 1876. I should note at this point that while I am interested in nineteenth-century baseball, I am by no means an expert on it, and so you should not be too surprised if I butcher the facts or make faulty assumptions, or call Cap Anson “Cap Anderson”. If you want a great historical presentation of old-time baseball, the best place to go is David Nemec’s The Great Encyclopedia of Nineteenth Century Major League Baseball. I believe that a revised edition of this book has been published recently, but I have the first edition. It is really a great book, similar in format to my favorite of the 20th century baseball encyclopedias, The Sports Encyclopedia: Baseball (or Neft/Cohen if you prefer). Like that work, only basic statistics are presented (no OPS+ or Pitching Runs, etc.), but you get the complete roster of each team each year, games by position, etc. And just like Neft/Cohen, there is a text summary of every season’s major stories, although Nemec writes these over the course of four or five pages, with pictures and trivial anecdotes, as opposed to the several paragraphs in the Neft/Cohen book. I wholeheartedly recommend the Nemec encyclopedia to anybody interested in the 19th century game.
That digression aside, the 1876 National League is still a different world then what we have today. The season is 60 games long, one team goes 9-56, pitchers are throwing from a box 45 feet away from the plate, it takes a zillion balls to draw a walk, overhand pitching is illegal, etc. But thankfully, you can make some sense of the statistics of this league, and while our tools don’t work as well, due to the competitive imbalance, the lack of important data that we have for later seasons, the shorter sample sizes as a result of a shorter season, etc., they can work to a level of precision that makes me comfortable to present their findings, with repeated caveats about how inaccurate they are compared to similar tools today. For the National Association, I could never reach that level of confidence.
What I intend to do over the course of this series is to look at the National League each season from 1876-1881. I chose 1881 for a couple reasons, the first being that during those seven seasons the NL had no other contenders to “major league” status (although many historians believe that other teams in other leagues would have been competitive with them--it's not like taking today’s Los Angeles Dodgers against the Vero Beach Dodgers). Also, in Bill James’ Historical Data Group Runs Created formulas, 1876-1881 is covered under one period (although 1882 and 1883 are included as well). That James found that he could put these seasons under one RC umbrella lead me to believe that the same could be done for BsR and a LW method as well. I will begin by looking at the runs created methodology here.
Run estimation is a little tricky as you go back in time. Unfortunately, there is no play-by-play database that we can use to determine empirical linear weights, and some important data is missing (SB and CS particularly). The biggest missing piece of the offensive puzzle though is reached base on error, which for simplicity’s sake I will just refer to as errors from hereon. In the 1880 NL, for instance, the fielding average was .901, and there were 8.67 fielding errors per game (for both teams). One hundred years later, the figures were .978 and 1.74. So you have something like five times as many errors being made as you do in the modern game.
When looking at modern statistics, you can ignore the error from an offensive perspective pretty safely. It will undoubtedly improve the accuracy of your run estimator if you can include it, but only very slightly, and the data is not widely available so we just ignore it, as we sometimes ignore sacrifice hits and hit batters and other minor events. But when there are as many errors as there were in the 1870s, you can’t ignore that. If you use a modern formula like ERP, and find the necessary multiplier, you will automatically inflate the value of all of the other events, because there has to be compensation somewhere for all of the runs being created as a result of errors.
So far as I know, there is only one published run estimator for this period. Bill James’ HDG-1 formula covers 1876-1883, and is figured as:
RC = (H + W)*(TB*1.2 + W*.26 + (AB-K)*.116)/(AB + W)
Bill decided to leave base runners as the modern estimate of H+W, and then try to account somewhat for errors by giving all balls in play extra advancement value. If you use the total offensive stats of the period to find the implicit linear weights, this is what you get:
RC = .730S + 1.066D + 1.402T + 1.739HR + .434W - .1081(AB - H - K) - .1406K
As you can see, the value of each event is inflated against our modern expectation of what they should be. I should note here that, of course, we don’t expect the 1870s weights to be the same as or even that similar to the modern weights. The coefficients do and should change as the game changes. That said, though, we have to be suspicious of a homer being valued at 1.74 runs and a triple at 1.40. The home run has a fairly constant value and it would take a very extreme context to lift its value so high. Scoring is high in this period (5.4 runs/game), but a lot of that logically has to be due to the extra errors. Three and a half extra errors per team game is like adding another 3.5 hits--it's going to be a factor in increased scoring.
To test RMSE for run estimators, I figured the error per (AB - H). I did this because I did not want the ever changing schedule length to unduly effect the RMSE. Of course, this does introduce the potential for problems because AB-H is much less a good proxy for outs in this period then it is today, as I will discuss shortly. I then multiplied the per out figure by 2153 (the average number of AB-H for a team in the 1876-1883 NL). In any case, doing this versus just taking the straight RMSE against actual runs scored did not make a big difference. Bill’s formula came in at 35.12 while the linearization was 30.65.
Of course what I wanted to do was figure out a Base Runs formula that worked for this period, as BsR is the most flexible and theoretically sound run estimator out there. What I decided to do was use Tango Tiger’s full modern formula and attempt to estimate some data that was missing and throw out other categories that would be much more difficult to estimate. I wound up estimating errors, sacrifice hits, wild pitches, and passed balls but throwing out steals, CS, intentional walks, hit batters, etc. Some of those events were subject to constantly changing rules and strategy (stolen bases and sacrifices were not initially a big part of the professional game) or didn’t even yet exist (Did teams issue intentional walks when it took 8 balls to give the batter first base? I am not a historian, but I doubt it. Hit batters did not result in a free pass until the 1887 in the NL). In the end, I came up with these estimates:
ERRORS: In modern baseball, approximately 65% of all errors result in a reached base on error for the offense. I (potentially dubiously) assumed that a similar percentage held in the 1870s, and used 70%. Then I simply figured x as 70% of the league fielding errors, per out in play (AB-H-K). x was allowed to be a different value for each season. Some may object to this as it hones in too much on the individual year and I certainly can understand such a position. However, the error rates were fluctuating during this period. In 1876 the league FA was .866; in 1877 it was up to .884; then .893, .892, .901, .905, .897, and .891. These differences are big enough to suggest that fundamental changes in the game may have been occurring from year-to-year.
James’ method had no such yearly correction, and if you force the BsR formula I will present later to use a constant x value of .134 (i.e. 13.4% of outs in play resulted in ROE), its RMSE will actually be around a run and a half higher then that of the linearization of RC. I still think that there are plenty of good reasons to use the BsR formula instead, but in the interests of intellectual honesty, I did not want to omit that fact.
It is entirely possible that a better estimate for errors could be found; there is no reason to assume that every batter is equally likely to reach on an error once they’ve made an out in play. In fact, I am sure that some smart mind could come along and come up with better estimates then I have in a number of different areas, and blow my formula right out of the water. I welcome further inquiry into this by others and look forward to my formula being annihilated. So don’t take any of this as a finished product or some kind of divine truth (not that you should with my other work either).
SACRIFICES: The first league to record sacrifices, so far as I can tell, was the American Association in 1883 and 1884. In those leagues, there was .0323 and .0327 SH per single, walk, and estimated ROE. So I assumed SH = .0325*(S + W + E) would be an acceptable estimate in the early NL. NOTE: Wow, did I screw the pooch on this one. The AA DID NOT track sacrifices in '83 and '84. I somehow misread the HB column as SH. We do no thave SH data until 1895 in the NL. So the discussion that follows is of questionable accuracy.
I did this some tie ago without thinking it through completely; in early baseball, innovations were still coming quickly, and it is possible that in the seven year interval, the sacrifice frequency changed wildly. George Wright recalled in 1915 (quoted in Bill James’ New Historical Baseball Abstract, pg. 10): “Batting was not done as scientifically in those days as now. The sacrifice hit was unthought of and the catcher was not required to have as good a throwing arm because no one had discovered the value of the stolen base.”
On the other hand, 1883 is pretty close to the end of our period, so while the frequency may well have increased over time, the estimate should at least be pretty good near the end of the line. One could also quibble with the choice of estimating sacrifices as a percentage of times on first base when, if sacrifices are not recorded, they are in actuality a subset of AB-H-K. Maybe an estimate based both on times on first and outs in play would work best. Again, there are a lot of judgment calls that go into constructing the formula, and so there are lots of areas for improvement.
WP and PB: These were kept by the NL, and there were .0355 WP per H+W-HR+E and .0775 PB per the same. So, the estimates are WP = .0355*(H + W - HR + E) and PB = .0775*(H + W - HR + E).
Then I simply plugged these estimates into Tango’s BsR formula. D of course was home runs, while A = H + W - HR + E + .08SH and C = AB - H - E + .92SH. The encouraging thing about this exercise was that the B factor only needed a multiplier of 1.087 (after including a penalty of .05 for outs) to predict the correct number of total runs scored. Ideally, if Base Runs was a perfect model of scoring (obviously it is not), we could use the same formula with any dataset, given all of the data, and not have to fudge the B component. The fact that we only had to fudge by 1.087 (compared to Bill James who to make his Basic RC work had to add walks into the B factor, take 120% of total bases, and add 11.6% of balls in play to B), could indicate that the BsR formula holds fairly well for this time when we add important, more common events like SH, errors, WP, and PB. Of course, perhaps Bill could get similar results using a more technical RC formula + estimation. The bottom line is, a fudge of only 1.087 will keep the linear weights fairly close to what we expect today. I don’t know for sure that they should be, but I’d rather error on the side of our expectations as opposed to a potentially quixotic quest to produce the lowest possible RMSE for a sample of sixty teams playing an average of 78 games each.
So the B formula is:
B = (.726S + 1.948D + 3.134T + 1.694HR + .052W + .799E + .727SH + 1.165WP + 1.174PB - .05(AB - H - E))*1.087
The RMSE of this formula by the standard above is 28.18. I got as low as 24.61 by increasing the outs weight to -.2, but I was not comfortable with the ramifications of this. As mentioned before, if one does not allow each year to have a unique ROE per OIP ratio, the RMSE is a much worse 32.20. Again, I feel a differently yearly factor is appropriate, but can certainly see if some feel this is an unfair advantage for this estimator when comparing it to others. The error of approximately 30 runs is a far cry from the errors around 23 in modern baseball, plus the season was shorter and the teams in this period averaged only 421 runs/season, so the raw number makes it seem smaller then it actually is. As I said before, you should always be aware of the inaccuracies when using any sabermetric method, but those caveats are even more important to keep in mind here.
Another way to consider the error is as a percentage of the runs scored by the team. This is figured as ABS(R-RC)/R. For sake of comparison, basic ERP, when used on all teams 1961-2002 (except 1981 and 1994), has an average absolute error of 2.7%. The BsR formula here, applied to all NL teams 1876-1883, has an AAE of 5.4%, twice that value. So once again I will stress that the methods used here are nowhere near as accurate as the similar methods used in our own time. Just for kicks, the largest error is a whopping 24.2% for the 1876 Cincinnati entry, which scored 238 runs but was projected to score 296. The best estimate is for Buffalo in 1882; they actually scored 500 versus a prediction of 501.
Before I move on too far, I have a little example that will illustrate the enormous effect of errors in this time and place. In modern baseball, there are pretty much exactly 27 outs per game, and approximately 25.2 of these are AB-H. We recognize, of course, that ROE in our own time are included in this batting out figure, and should not be, but any distortion is small and can basically be ignored.
Picking a random year, in the 1879 NL, we know that there were 27.09 outs/game since we have the innings pitched figure. How many batting outs were there per game? Well, if the modern rule of thumb held, there should be just about 25.2. There were 28.01. So there are more batting outs per game then there are total outs in the game. With our error estimate subtracted (so that batting outs = AB - H - E), we estimate 24.60. Now this may well be too low, or just right, or what have you. Maybe I it should have been 50% of errors put a runner on first base instead of 70%. I don’t know. What I do know is that if you pretend errors do not exist, you are going to throw all of your measures for this time and place out of whack. Errors were too big of a factor in the game to just be ignored as we can do today.
Let’s take a look at the linear values produced by the Base Runs formula, as applied to the entire period:
Runs = .551S + .843D + 1.126T + 1.404HR + .390W + .569E + .081SH + .280PB + .278WP - .145(AB - H - E)
This is why I felt much more comfortable with the BsR formula I chose, despite the fact that there were versions with better accuracy. These weights would not be completely off-base if we found them for modern baseball. Whether or not they are the best weights for 1876-1883, we will have to wait for when brighter minds tackle the problem or when PBP data is available and we can empirically see what they are. But to me, it is preferable to accept greater error in team seasonal data but keep our common sense knowledge of what events are worth rather then to chase greater accuracy but distort the weights.
This is still not the formula that I am going to apply to players, though. For that, I will use the linear version for that particular season. Additionally, for players, SH, PB, and WP will be broken back down into their components. What I mean is that we estimate that a SH is worth .081 runs, and we estimated that there are .0325 SH for every S, W, and E. .081*.0325 = .0026, and therefore, for every single, walk, and error we’ll add an additional .0026 runs. So a single will be worth .551+.0026 = .554 runs. We’ll also distribute the PB and WP in a similar way.
There are some drawbacks to doing it this way. If Ross Barnes hits 100 singles, his team may in fact lay down 3.25 more sacrifices. But it will be his teammates doing the sacrificing, not him. And we would assume that good hitters would sacrifice less then poor hitters, and this method assumes they are all doing it equally.
On the other hand, though, we are just doing something similar in spirit to what a theoretical team approach does--crediting the change in the team’s stats as a direct result of the player to the player. Besides, there’s really no other fair way to do it (we don’t want to get into estimating SH as a function of individual stats, and even if we did, we have no individual SH data for this period to test against). Also, in the end, the extra weight added to each event will be fairly small, and I am much more comfortable doing it with the battery errors which should be fairly randomly distributed with regards to which particular player is on base when they occur.
Then there is the matter of the error. Since the error is done solely as a function of AB-H-K, we could redistribute it, and come up with a different value for a non-K out and a K out, and write errors out of the formula, and have a mathematically equivalent result. However, I am not going to do this because I believe that, as covered previously, errors are such an important part of this game that we should recognize them, and maybe even include them in On Base Average (I have not in my presentation here, but I wouldn’t object if someone did) in order to remember that they are there. I think that keeping errors in the formula gives a truer picture of the linear weight value of each event as well, as it allows us to remember that the error is worth a certain number of runs and that outs, actual outs, have a particular negative value. Hiding this by lowering the value of an out seems to erase information to me.
I mentioned earlier that each year will have a different x to estimate errors in the formula x(AB-H-K). They are: 1876 = .1531, 1877 = .1407, 1878 = .1368, 1879 = .1345, 1880 = .1256, 1881 = .1184.
At this point, let me present the weights for the league as a whole in each year in 1876-1881, and then the ones with SH, PB, and WP stripped out and reapportioned across the other events. The first set is presented as (S, D, T, HR, W, E, AB-H-E, SH, PB, WP). The second is presented as (S, D, T, HR, W, E, AB-H-E).
1876: .552, .853, 1.146, 1.417, .386, .570, -.147, .085, .289, .287
1876: .588, .886, 1.178, 1.417, .422, .606, -.147
1877: .563, .862, 1.152, 1.414, .398, .581, -.153, .079, .287, .285
1877: .598, .894, 1.184, 1.414, .433, .616, -.153
1878: .546, .846, 1.138, 1.417, .380, .564, -.144, .087, .289, .287
1878: .581, .879, 1.171, 1.417, .415, .599, -.144
1879: .543, .830, 1.108, 1.397, .385, .560, -.140, .082, .275, .273
1879: .577, .861, 1.139, 1.397, .419, .594, -.140
1880: .537, .825, 1.105, 1.400, .378, .554, -.137, .086, .277, .275
1880: .571, .856, 1.136, 1.400, .412, .588, -.137
1881: .560, .859, 1.149, 1.415, .395, .578, -.151, .080, .287, .285
1881: .595, .891, 1.182, 1.415, .430, .613, -.151
Next installment, I’ll talk a little bit about replacement level, the defensive spectrum, and park factors.