Wednesday, May 26, 2021

Rate Stat Series, pt. 4: Players as Teams

 A dynamic run estimator is a run estimator that allows offensive events to interact with each other, such that the value of a given event is not fixed as would be the case in a linear weights formula (e.g. a single is worth .50 runs), but rather is dependent upon all of the other components of the batting line. Dynamic run estimators are great in theory, since the run scoring process for a team is obviously dynamic and not linear. However, there are two issues:

1. They are harder to design than linear estimators. Any idiot with a spreadsheet and a dataset can run a linear regression on runs scored and have a linear estimator when they are done. It may not be a good one, but it will be functional and will probably have a low RMSE when estimating team runs scored. To develop a dynamic model, one must consider the run scoring process and produce a simplified model, but not so simplified as to not produce reasonably accurate estimates.

This is not a series about run estimators, but the most commonly used dynamic run estimator, Bill James’ Runs Created, suffers from flaws that make it unable to handle extreme offenses. A much better model, David Smyth’s Base Runs, is powerful and will be used here.

2. They are not appropriate to apply to individual offensive statistics. Dynamic estimators always involve multiplying base runners by some factor representing advancement of baserunners (in Runs Created that’s the end of the story, Base Runs accounts for the unique nature of home runs). This multiplication is inappropriate when applied to an individual player, as now Frank Thomas’ high OBA is multiplied directly with his high power which advances runners. In reality, there is some interaction, but Thomas’ impact is diluted by being just 1/9th of the lineup. Inputting his statistics into a dynamic run estimator produces an estimate of how many runs a team would score if each batter hit like Thomas.

Due to this issue, I do not advocate applying dynamic run estimators directly to individuals, but this post will still address the rate stat implications of such applications. Later we will discuss theoretical team methods that allow the use of a dynamic run estimator while still accounting for the fact that the player is just one of nine in the lineup.

This series will now discuss what I believe to the be the proper rate stats for a particular framework for evaluating individual offense. One of my objectives is that for each option of a framework for building a rate stat presented, there be at least one variation that is linearly comparable and one that is ratio comparable. I’ve defined those terms as I use them at length before, so here I will be brief:

* A statistic is linearly comparable if the difference between two figures is meaningful. A hitter with a .400 OBA would reach base 100 times more than a hitter with a .300 OBA over 1000 PA.

* A statistic is ratio comparable if the ratio between two figures is meaningful. Our .400 OBA player reached base 33.3% more frequently than the .300 OBA player

Ideally, our metric will facilitate both types of comparison, but if not, I will endeavor to present an alternative formulation that fills the gap. I will not propose any metrics that are neither linearly comparable or ratio comparable because they are the scourge of sabermetrics (hello OPS).

The underlying principle of the discussion that follows for the three frameworks (treating the player as a team, a full linear model, and a theoretical team model) is that the rate stat should be consistent with the run estimator used. If the run estimator treats the player as if he is a team, then the corresponding rate stat should treat the player as if he is a team.

In this case, that makes it very simple. The proper denominator for a team rate stat is outs. If you apply Runs Created, Base Runs, or some other run estimator directly to an individual player, the proper denominator is outs.

At this point in the discussion, this may ring as a somewhat hollow declaration, as I have only indirectly made the case for why we might want to use a denominator other than outs for an individual when it is so clearly the proper choice for a team. Since I’m suggesting that outs are the proper choice for this framework, I’ll defer that case for later.

In this case, I advocate for using outs when applying a dynamic run estimator to a team because it is the only consistent treatment. The only justification for going down this path (other than needing something quick and dirty) is a theoretical exercise – how many runs would a team that hit like Frank Thomas score? While I don’t think this theoretical result is appropriate for attempting to value Thomas’ contribution the 1994 White Sox, it at least does have an interpretation. If you start mixing frameworks, you really have a mess on your hands. There’s no good reason (other than crude estimation) to apply a dynamic run estimator directly to an individual; there’s no sense in deviating from the corresponding rate stat in order to try to make the results more comparable to a better approach to evaluating individual offensive contribution. Just use the better approach, and if you insist on misapplying a dynamic run estimator to individual players, make outs the denominator so that at least you have a theoretically coherent suite of metrics.

I should note that Bill James in the 1980s took this entire process to its logical conclusion. After applying Runs Created to individuals, dividing by outs, and multiplying by a constant that was close to the league outs/game for the definition of outs chosen, he went a step further and used the Pythagorean theorem to estimate the winning percentage that this team would have if it allowed an average number of runs. He then converted it to wins and losses by using the number of outs the player made to define games, which caused all kinds of problems, but at least he was committed.

This will be the first of several times that I’ll run a leaderboard for the 1994 AL using a particular framework. Here we have the top 5 and bottom 5 performers with at least 200 PA in Base Runs/Out. RAA is “Runs Above Average” and is calculated simply as (BsR/O – LgR/O) * Outs. Spoiler alert: No matter how we slice it, Frank Thomas is going to come out as the leading hitter in this league, as he raked .353/.492/.729 on his way to a second consecutive MVP award.


I am showing at least one more decimal place on each metric than I usually would just to allow for a little more precise calculation if you’re following along; it is no way a statement about the significance of the ten-thousandths of runs per out. 

Runs per out can of course be scaled; Bill James multiplied it by the league average outs/game appropriate given the categories be considered in the computation of outs. For instance, in this case, since we’re defining outs as AB – H, the average outs/game will be around 25.2 (for the 1994 AL it was 25.19). A more complete accounting of outs, like AB – H + CS + SH + SF + DP, would get close to 27 outs/game. While putting individual contribution on a team games basis is nonsensical on some level, since it is just a scalar multiplier it causes no real distortion and provides a scale that is easily understandable, in the same manner that ERA or K/9 are understood by everyone other than Matt Underwood and Harold Reynolds. 

Wednesday, May 12, 2021

Rate Stat Series, pt. 3: Teams

If I tell you that three teams in the same league-season played the same number of games (113), and that one of them scored 679 runs, another scored 670, and the third scored 633, how confident would you be in using this limited data to rank the productivity of their offenses? As usual in this series, we are ignoring park factors and other contextual factors (like quality of opposition/not having to face one’s own pitching staff); since they are from the same league-season, you don’t need to worry about whether the win value of each team’s runs was the same. Assume also that runs will be distributed across games by a known distribution like Enby, so the distribution is also not a differentiator. Assume that we don’t care about any “luck”; the actual total is what matters, not what a run estimator came up with. What else do you need to know?

I would contend that given the (admittedly restrictive) parameters I’ve placed on the exercise, you now know almost everything you need to know. In a small number of cases, and to a small extent, you are missing valuable information – but for most situations, you should need no additional information.

Now suppose I told you something similar about three players: same league season, same number of games played (111), and three runs created estimates: one player created 106 runs, one 92, and one 88. Do you feel like you need any additional information to put these players in the proper order of offensive productivity?

I hope that your answer here is yes, and a lot of it. I’ve told you how many games each have played, but that doesn’t tell you how many opportunities they’ve had at the plate. Sure enough, in this case one of the players had substantially fewer plate appearances than the others (489, 490, 451 respectively). Given that the player who created 90 runs had 39 more plate appearances than the player who created 86, it seems likely that the latter player was actually more productive on a rate basis.

I did not tell you how many plate appearances each of the three teams had in their 113 games; I don’t think it’s relevant to the question at hand, but the answer is 4493, 4611, and 4556 respectively. Why do we need to know plate appearances (or something) in the case of players, but not in the case of teams? Understanding this gets to the heart of the reason this series needs to exist at all, why applying the same rate stat to team offenses and player offense may not work as intended.

In the previous installment, I asked the question: “Where do plate appearances come from?” The answer is that every inning (excluding walkoff situations) starts with three PAs guaranteed, and only by avoiding outs (reaching base and not being subsequently retired on the bases) can a team generate additional plate appearances.

From a team perspective, then, plate appearances are not an appropriate denominator for a rate stat, because differences in team plate appearances are the result of differences in performance between the teams. To return to the three teams discussed above, they are the 1994 Indians, Yankees, and White Sox respectively. The Indians had the fewest PA of the three yet scored the most runs. Does this mean that their offense, which already scored more runs than the other two clubs, was even more superior than the raw numbers would suggest?

An offense does not set out to maximize its plate appearances, nor does it set out to score the maximum number of runs it can in the minimum number of plate appearances. An offense sets out to maximize its total runs scored. Plate appearances are a function of the rate at which a team makes outs. At this point it might be helpful to consider the three teams:



New York’s OBA was 22 points higher than Cleveland’s and thus they generated an extra plate appearance per game. When ranking team offenses, it wouldn’t make sense to penalize the Yankees for this, which would be the case if we used R/PA. The difference in plate appearances simply reflects the different manner in which New York and Cleveland went about creating runs. For a team, plate appearances are inextricably linked with their OBA. Each inning, a team attempts to score as many runs as it possibly can before making three outs. It’s possible to score one run in a complete inning with as few as four or as many as seven plate appearances. Whether a team uses four, five, six, or seven plate appearances to score a single run is irrelevant in terms of that run’s impact on them winning or losing the game (*). Thus outs or an equivalent like innings are the correct choice for the denominator of a team rate stat.

(*) I am speaking here simply about the direct impact of the runs scored and not any downstream effects or the predictive value of team performance. Perhaps the team that uses seven PA to score one run benefits by wearing down the opposing pitcher or is more likely to have success in the future because they had four of seven batters reach base compared to one in four for the team that only needed four PA. Here we’re just focused on the win value directly attributable to the run scored and not any secondary or predictive effects.

The fact that outs are fixed for each team each inning (ignoring walkoffs) means that outs are also fixed for each team each game (ignoring walkoffs, rainouts, extra innings, and foregone bottom of the ninths). Which means that outs are also fixed for each team each season (ignoring those factors and cases in which teams don’t play out their full schedules, or have to play tiebreakers), which means that R/G and raw seasonal runs scored total are essentially equivalent to looking at R/O for a team. So for the question I asked at the beginning of the article, just knowing that the three teams had played an equal number of games, we had a pretty good idea how they would “truly” rank using R/O.

For players, this is not at all the case, since even in an equal number of games, players will get different numbers of plate appearances for a variety of reason (batting order position, the team’s OBA (remember, higher OBA teams will generate more PA), whether or not they play the full game), a fact that is intuitive to most baseball fans. What is less intuitive, though, is that even in the same number of plate appearances, players can make very different numbers of outs. Since we’ve already accepted that team OBA defines how many plate appearances a team will generate, it isn’t much of a leap to conclude that if we have two players who create the same number of runs (using a formula that doesn’t explicitly account for their impact on the team’s OBA) in the same number of plate appearances, the player who makes fewer outs was more productive when we consider the totality of their offensive contribution. Even though the two players were equally productive in their plate appearances, the player who made fewer outs generated more plate appearances for his teammates, a second-order effect that needs to be considered when evaluating individual offensive contribution. For teams, the runs scored total already reflects this effect.

This would be an appropriate time to note that this series is focused on evaluating offenses, but of course every offensive metric can be reviewed in reverse as a defensive metric. However, since the obvious denominator for teams is outs, it is also the obvious denominator for individual pitchers. We don’t need to worry about a pitcher’s impact on his team’s plate appearances – when he is in the game, he is solely responsible (setting aside the question of how the team’s performance should be allocated between the pitcher and his fielders) for the number of plate appearances the opponent generates, and his goal is to record three outs while minimizing the number of runs he allows, regardless of how many opponents come to the plate. Outs are clearly the correct denominator for the rate stat, and innings pitched are nothing more than outs/3 (and even better, IP account for all outs, including many that don’t show up in the standard statistical categories).

In thinking about the development of early baseball statistics and the legacy of those standard statistics on how the overwhelming majority of fans thought about baseball before the sabermetric revolution took hold, it is striking that the early statisticians understood these concepts as they applied to pitchers. When pitchers were completing almost all their starts, simple averages of earned runs allowed sufficed, for the same reason that team R/G tells you most everything you need to do. As complete games became rarer, ERA took hold, properly using innings in the denominator. For most of the twentieth century, and even post-sabermetric revolution, baseball fans are conditioned to think about innings pitched as the denominator for all manner of pitching metrics – even those like strikeout and walk frequency for which plate appearances would make a much more logical denominator. (Of course, present day sabermetrics has embraced metrics like K% and W% for pitchers, but the per inning versions remain in use as well).

The parallel development of offensive statistics resulted in the opposite phenomenon. While early box scores tracked “hands out” (essentially outs made) for individual batters, batting average eventually became the dominant statistic. Setting aside the issues with “at bats” and how they distort people’s thinking and saddled us with the mouthful of “plate appearances” to describe the more fundamental quantity of the two, the standard batting statistics have conditioned fans to think about batting rates (walk rate, home run rate, etc.) in the correct manner (or one adjacent to being correct, depending on whether at bats or plate appearances are the denominator), but leave people struggling with how to properly express a batter’s overall productivity. Again, this is the opposite problem of how pitching statistics were traditionally constructed. One can imagine that it all might be very different had the Batting Average taken the form of a hit/out ratio rather than hits/at bats.

Wednesday, April 28, 2021

Rate Stat Series, pt. 2: PA Generation

This is a little bit of a detour and certainly nothing new (I don’t know who originally laid out this logic/math – the earliest use I’m aware of was in 1960 by D’Esopo & Lefkowitz as part of their Scoring Index model), but I think a discussion of it is appropriate in the context of this series, and I will later make use of these formulas.  It’s also ground I covered in the original series, but I think my explanation this time is slightly more coherent.

Each batting team starts each inning (excluding scenarios where a walkoff is possible) with three plate appearances guaranteed. Thus each team starts each game with twenty-seven plate appearances guaranteed (excluding scenarios where the home team forgoes batting the bottom of the ninth, rainouts, post-2020 doubleheaders, etc.). Any plate appearances beyond that must be earned by batters avoiding outs. Since it’s more natural to think of a positive outcome rather than the avoidance of a negative outcome, I will simplify and say that each extra plate appearance must be earned by a batter reaching base (and not being subsequently retired on the basepaths).

For the sake of discussion (and keeping with the simple set of statistics being used in the metrics in this series), I’m going to ignore the existence of baserunning outs, including caught stealing, pickoffs, outs stretching, outs advancing, and runners retired on double/triple plays (although not on fielder’s choices, since the batter is charged with an out in that case). I’m going to assume that the out rate is the complement of on base average, which in this series will be defined simply as (H + W)/(AB + W). In reality, considering all the ways in which outs can be made, it would be a more involved equation (I’ve used the acronym NOA for Not Out Average and OA for the complement, Out Average) which would look something like this, although it still doesn’t think I’ve accounted for every possible event (you try incorporating fielders’ choices without complicating the equation significantly):

NOA = (H + W + HB + CI + ROE – CS – DP – Outs Stretching – Outs Advancing – Pickoffs – 2*TP)/(AB + W + HB + SF + SH + CI)

Alternatively, for a team when LOB data is available (and ignoring the walkoff situation), you could have OA = (Plate Appearances – Runs Scored – Left On Base)/Plate Appearances. All of this is just an attempt to calculate, as best we can from the available statistics we have restricted ourselves to, Outs/Plate Appearances. NOA or OA as appropriate could be substituted for OBA in the equations that follow as long as the appropriate corresponding adjustments are made to the numerator.

Let’s assume for the purpose of developing an equation for team plate appearances that the OBA is constant across each of the nine batters in the lineup and doesn’t vary for any other reason (this is obviously never true, but it is a fine simplifying assumption for modeling PA generation). Then a team will start an inning with three plate appearances. For each of those three guaranteed PAs, there is a probability (equal to OBA, given our assumption) that the batter avoids an out (reaches base, given that there are no baserunning outs). This increases the expected number of plate appearances by OBA.

It doesn’t stop there, though. Each additional PA that is generated also has an OBA chance of creating an additional PA, which itself has an OBA chance of creating an additional PA. Thus, for each of the guaranteed PA, the expected final number of team PA is:

OBA + OBA*OBA + OBA*OBA*OBA + … = OBA + OBA^2 + OBA^3 + … OBA^n

which when n is infinity and OBA is between 0 and 1 (which it must be by definition) resolves to:

OBA/(1 – OBA)

The 1994 AL had an OBA of .343. Thus, each guaranteed plate appearance should have generated .343/(1 - .343) = .522 additional plate appearances. In an average inning, starting with three guaranteed PA, we would expect 3 + 3*.522 = 3*(1 + .522) = 4.566 PA, and thus in a game we would expect 9*4.566 = 41.09 PA. Note that instead of calculating the .522 additional PA, we can simplify this to 3/(1 – OBA) for an inning or 27/(1 – OBA) for a game. In reality there were 39.24 PA, so we have an unacceptable 4.7% error. What went wrong?

I’m mixing definitions of plate appearances and definitions of OBA incorrectly, and also ignored that the three guaranteed PA are equal to the number of outs permitted in the inning. In order to estimate the number of plate appearances per inning or game consistently, we need to divide the average number of outs/game by 1 – OBA:

PA/G = (O/G)/(1 – OBA)

The definition of outs that corresponds to our simple (H + W)/(AB + W) complement of out average is AB – H. In the 1994 AL there were 25.19 outs/game using this definition, so our expected PA/G is:

25.19/(1 - .353) = 38.34

The actual average was 38.35; we’re off due to rounding as this is now just a mathematical truism since by our simplified definitions plate appearances = outs + times on base. Using this equation to estimate team PA/G from their OBA for the 1994 AL, the RMSE is .259, which is about .7% of the average PA/G. We shouldn’t expect perfect accuracy at the team level since team PA will be affected by different quantities of all the statistical categories we’re ignoring that have an impact on the actual number of PA a team generates, as well as differences in number of extra inning games, foregone bottom of the ninths, and walkoff-shortened innings.

The key points to keep in mind as we move forward in discussing rate stats are:

1.      The number of plate appearances a team will get is a function of their out rate, and simplifying terms we can very accurately estimate team PA as a function of on base average

2.      Since players have an impact on the number of plate appearances their team gets, and thus the number of plate appearances they get, a proper rate stat for measuring overall offensive productivity must account for that impact

Thursday, April 15, 2021

Almost Perfect

In my earlier days as a baseball fan, I was really interested in no-hitters, and outside of the Indians winning the World Series, my most fervent desire as a fan was to witness one even if only on the radio. Eventually this faded, due to some combination of growing jaded about the extent to which baseball fans sometimes elevate trivial events above game outcomes, the pernicious influence of Voros McCracken on how I thought about the hits column for pitchers, and after fifteen years of intense baseball-watching finally witnessing one (I'm now up to five).

Perfect games retain a bit more of their mystique for me, due to being much more rare (someone who has watched as many games over the years as I have is bound to have seen a no-hitter, but one can't really expect to see a perfect game) and not relying on any arbitrary distinction between hits and errors (which of course doesn't affect all no-hitters). The three closest games I have taken in to being perfect games prior to last night were Mike Mussina against the Indians in 1997 and Armando Galarraga's should-have been perfect game against the Indians in 2010. The latter game is case in point of what I meant about fans sometimes being more interested in trivial events than game outcomes - there was more outcry in favor of replay as a result of that game then there was cumulatively from many calls that much more directly influenced which team won a given game.

Last night's effort by Carlos Rodon combined elements of both of the ninth innings of these games in the way that people who believe in hocus pocus should embrace. From Galarraga's, we took the extremely close play at first base, with Josh Naylor playing the role of Jason Donald, desperately trying to reach first after making weak contract towards first base. In this case, the play was actually much closer, but no replay was required as the call on the field was that Jose Abreu beat him to the bag by a narrow margin. 

From the Mussina game, we borrowed the man, lineup slot, and fielding position to break it up. With one out in the ninth, the Indians catcher. Sandy Alomar singled off Mussina, while Roberto Perez was only hit in the back foot with a slider, but history repeated itself in who ended it. Of course, if Rodon had to lose the perfect game, he got the better outcome than the other two, as he at least got to keep the no-hitter.

Naturally, all of the near perfect games I've seen have been pitched against the Indians. In addition to the infinitely more important distinction of now having the longest World Series drought, after Joe Musgrove's no-hitter for the Padres, the Indians now have the longest drought between no-hitters, it having been nearly forty years since Len Barker's perfect game.

I was keeping score of the Mussina game and Rodon's effort last night, but not the Galarraga game, which I listened to on the radio while I watched some other game on TV. 



Wednesday, April 14, 2021

Rate Stat Series, pt. 1: Introduction

This blog has existed for sixteen years now, and yet with the exception of some (relatively) recent stuff I’ve written about the Enby distribution for team runs per game and the Cigol approach to estimating team winning percentage from Enby, almost all of the interesting sabermetric work appeared in the blog’s first five years, and most in the first year or two.

There are a number of reasons for that - one is that when I started, I was a college student with a lot more free time on his hands than I have with a 9-5. Related, I was also more eager to spend a lot of time staring at numbers on my free time when I didn’t spend a good portion of my day staring at numbers. Remember the Bill James line about how a column of numbers that would put an actuary to sleep can be made to dance if you put Bombo Rivera’s picture on the flip side of the card? Sometimes the numbers do indeed dance, but the actuary in question would rather watch a ballgame or read about the Battle of Gravelines than manipulate them in the evening, dancing or no.

More generally, there has been much less to investigate in the area of sabermetrics that I primarily practice, which I will call for the lack of a better term “classical sabermetrics”. I would define classical sabermetrics as sabermetric study which is primarily focused on game-level (or higher, e.g. season, player career, etc.) data that relates to baseball outcomes on the field (e.g. hits, walks, runs scored, wins). Classical sabermetrics is/was the primary field of inquiry of those I have previously called first or second-generation sabermetricians.

Classical sabermetrics is not dead, but to date the last great achievement of the field was turned in by Voros McCracken when he developed DIPS. I’m not arrogant enough to declare that nothing more will ever be found in the classical field, and there is still much work to be done, but at least as far as I can see, it is highly likely that it will consist of tinkering and incrementally improving work that has already been done, and probably with little impact on the practical implementation of sabermetric ideas. For example, I still would love to find a modification to Pythagenpat that works better for 2 RPG environments, or a different run estimator construct that would preserve the good properties of Base Runs while better handling teams that hit tons of triples. All of this is quite theoretical, and of no practical value to someone who is attempting to run the Pirates.

Which increasingly is what sabermetric practitioners are attempting to do, whether directly through employment by major league teams, or indirectly through publishing post-classical sabermetric research in the public sphere. Let me be very clear: this is not in any way a lament for a simpler, purer time in the past. I think it’s wonderful that sabermetric analysis has transcended the constraints of the data used in its classical practice and is exerting an influence on the game on the field.

Notwithstanding, I am still a classical sabermetrician, not because I don’t value the insight provided by post-classical sabermetrics but because I don’t have some combination of the skillset or the way of thinking or the resources or the drive to become proficient enough in newer techniques to offer anything of value in that space. Thus it is natural that I have less to share here.

The topic that I am embarking on discussing is squarely in the realm of “quite theoretical and of no practical to someone who is attempting to run the Pirates”. About fifteen years ago, I started writing a “Rate Stat Series”, and aborted it somewhere in the middle. I have stated several times that I intend to revisit it, but until now have not. The Rate Stat Series was and now is intended to be a discussion of how best to express a batter’s overall productivity in a single rate stat. I should note three things that it is not:

1. The discussion is strictly limited to the construction of a rate stat measuring overall offensive productivity, not a subset thereof. I am not suggesting that if you are measuring a batter’s walk rate, strikeout rate, ground-rule double rate, or any other component rate you can dream up, that you should follow the conclusions here. For most general applications, plate appearances makes perfect sense as the denominator for a rate for any of those quantities. There may be reasons to follow a sort of decision tree approach that results in different denominators for some applications (McCracken was an innovator in this approach, in DIPS and park factors). All of that is well and good and completely outside the scope of this series.

2. The premise presupposes that the unit of measurement of a batter’s productivity has already been converted to a run-basis. Thus it is not a question of OPS v. OTS v. OPS+ v. 1.8*OBA + SLG v. wOBA v. EqA v. TAv v. whatever, but rather what the denominator for a batter’s estimated run contribution should be. The obvious choices are outs and plate appearances, but there are other possibilities. Spoiler alert: My answer is “it depends”.

3. Revolutionary, groundbreaking, or any other similar adjective. I’m attempting to describe my thoughts on methods that already exist and were created by other people in a coherent, unified format.

In sitting down to write this, I realized I made two fundamental mistakes in my first attempt:

1. I was attempting to “prove” my preferences mathematically, which is not a bad thing in theory, but some of what I was doing begged the question and some of this discussion is of a theoretical nature that lends itself more to logical reasoning/“proofs” than to mathematical “proofs”. I’ve tried to anchor my conclusions in math, logic, and reason where possible, but have also embraced that some of it is subjective and must be so.

2. I posted pieces before I finished writing the whole thing, or even knowing exactly where it was going.

These are rectified in this attempt – all of my assertions are wildly unsupported and as I hit post, all planned installments exist in at least a detailed outline form. While I have attempted to avoid the two mistakes I identified in the previous series, as I look at this series in full I can see I have may have just replaced them with two characteristics that will make reading this a real chore:

1. I’m overly wordy; repeating myself a lot and trying to be way too precise in my language (although I fear not as precise as the topic demands). There’s a lot of jargon in an attempt to delineate between the various concepts and methodological choices.

2. There’s way too much algebra; where possible, I didn’t want to just assert that mathematical operations resolved in a certain way and give an empirical example that backs me up, so there’s a lot of “proofs” that will be of no general interest.

Allow me to close by laying some groundwork for future posts. I am going to use the 1994 AL as a reference point, and when I use examples they will generally be drawn from this league-season. Why have I chosen the 1994 AL?

1. 1994 was the year I became a baseball fan, and I was primarily focused on the AL at that time, so it is nostalgic. I have not turned into a get off my lawn type who thinks that baseball reached its zenith in 1994 and it’s all been downhill since, but I do think that about 1994 Topps, the greatest baseball card set of all-time.

2. As the year in which the “silly ball era” really broke out, and due to the strike shortening the season, there are some fairly extreme performances that are useful when talking about the differences between rate stat approaches.

As discussed, this series starts from the premise that a batter’s contribution is measured in terms of runs, and work from there. This approach does not require the use of any particular run estimator, although one of my assertions is that the choice of run estimator and the choice of rate/denominator for the rate are logically linked. There are three types of run estimators that I will use in the series: a dynamic model, a linear model, and a hybrid theoretical team model.

In order to avoid differences in the run estimator(s) used unduly influencing differences in the resulting rate stats, I am going to anchor a set of internally consistent run estimators in the reference period of the 1994 AL. It will come as no surprise if you’ve read anything I’ve written about run estimators in the past that I am using Base Runs for this job. The point of this series is not to tell you which particular run estimator to use or how to construct it. It really doesn’t matter which version of Base Runs I use (if you are still stuck on Runs Created, there’s no judgment from this corner, at least for the duration of this discussion), or which categories I include in the formula – this is about the conceptual issues regarding the rate that you calculate after estimating the batter’s run contribution, so I am keeping it very simple, looking just at hits, walks, and at bats (thus defining outs as at bats minus hits) and ignoring steals/caught stealing, hit batters, intentional walks, sacrifices, etc..  Since I’m doing this with the run estimator, I will also do it with most other statistics I cite – for example, throughout this series OBA will be (H + W)/(AB + W), and PA will just be AB + W.

A version of Base Runs I have used is below. It’s not perfect by any means; it overvalues extra base hits as we’ll see below, but again, the specific estimator is for example only in this series – the thinking behind constructing the resulting rates is what we’re after:

A = H + W – HR

B = (2TB - H – 4HR + .05W)*.78

C = AB – H

D = HR

BsR = (A*B)/(B + C) + D

Typically, any reconciliation of Base Runs to a desired estimate number of runs scored for an entity like a league is done using the B factor, since it is already something of a balancing factor in the formula, representing the somewhat nebulous concept of “advancement” while the other components (A = baserunners, C = outs, D = automatic runs) represent much more tightly defined quantities. In order to force the Base Runs estimate for the 1994 AL to equal the actual number of runs scored, you need to replace the .78 multiplier with .79776, which can be determined by first calculating the needed B value (where R is the actual runs scored total):

Needed B = (R – D)*C/(A – R + D)

Divide this by (2TB – H – 4HR + .05W) and you get a .79776 multiplier. I usually don’t force the estimated runs equal to the actual runs, but for this series, I want to be internally consistent between all of the estimators and also be able to write formulas using league runs rather than having to worry about any discrepancies between league runs and estimated runs.

So our dynamic run estimator (BsR) used throughout this series will be:

A = H + W – HR = S + D + T + W

B = (2TB - H – 4HR + .05W)*.79776 = .7978S + 2.3933D + 3.9888T + 2.3933HR + .0399W

C = AB – H = Outs

D = HR

BsR = (A*B)/(B + C) + D

To be consistent, I will also use the intrinsic linear weights for the 1994 AL that are derived from this BsR equation as the linear weights run estimator. The intrinsic linear weights are derived through partial differentiation of BsR with respect to each component. If we define A, B, C, and D to be the league totals of those, and a, b, c, and d to be the coefficient for a given event in each of the A, B, C, and D factors respectively, than the linear weight of a given event is calculated as:

LW = ((B + C)*(A*b + B*a) – A*B*(b + c))/(B + C)^2 + d

For the 1994 AL, this results in the equation, where RC is to denote absolute runs created:

LW_RC = .5069S + .8382D + 1.1695T + 1.4970HR + .3495W - .1076(outs)

We will also need a version of LW expressed in the classic Pete Palmer style to produce runs above average rather than absolute runs. That’s just a simple algebra problem to solve for the out value needed to bring the league total to zero, which results in:

LW_RAA = .5069S + .8382D + 1.1695T + 1.4970HR + .3495W - .3150(outs)

I am ignoring any questions about what the appropriate baseline for valuing individual offensive performance is. Regardless of where you side between replacement level, average, and other less common approaches, I hope you will agree that average is a good starting point which can usually be converted to an alternative baseline much more easily than if you start with an alternative baseline. Average is also the natural starting point for linear weights analysis since the empirical technique of calculating linear weights based on average changes in average run expectancy is by definition going to produce an estimate of runs above average.

Later we will also have some “theoretical team” run estimators built off this same foundation, but discussion of them will fit better when discussing that concept in greater detail.

I will also be ignoring park factors and the question of context in this series (at least until the very end, where I will circle back to context). Since I am narrowly focused on the construction of the final rate stat, rather than a full-blown implementation of a rating system for players, park factors can be ignored. Since I am anchoring everything in the 1994 AL, the context of the league run environment can also be ignored since it will be equal for all players once we ignore park factors.

Thursday, April 01, 2021

Give Us This Day Our Daily Ball

Rob Manfred, who art Commissioner

Halloweth be our game

Thy rule changes be undone, thy no longer assault fun

In 2022 as it was in 2002

Give us this day our daily ball

And reconcile with Tony Clark as we reconcile to runners on in extra innings

And lead us not into strike or lockout

And deliver us from pitchers hitting

For thine is the office and the power and the responsibility until 2024

Play ball

Tuesday, March 30, 2021

2021 Predictions

I’m not telling you anything you don’t already know, but the 2021 season will be the hardest to forecast in recent times. While there weren’t people doing systematic forecasts as we would recognize them in the sabermetric era at the time, the last season that I believe would have posed a greater challenge to forecasters was the 1946 season in which so many players returned from military service. 2021 is hard to predict because we only had a sixty-game season on which to judge player’s current performance, and because there was no minor league season at all. The only season that would have been tougher to predict would have been 1995, if any poor sap had attempted that using the rosters as they stood prior to the labor ceasefire.

Of course, this does not pose any particular challenge to me in writing this, because I don’t do a systematic forecast of my unknown. I usually use publicly available player projections as a starting point, making my own seat of the pants adjustments for performance and playing time; because of the additional inaccuracy inherent to such an exercise in trying to predict 2021, I have eschewed that and just used team-level projections as a starting point. Since this is an exercise in fun (baseball is supposed to be fun) and not a serious sabermetric endeavor, cutting out the trappings of formal analysis will not harm it – you can’t go down any lower.

AL East

1. New York

2. Tampa Bay (wildcard)

3. Toronto

4. Boston

5. Baltimore

I’ve been cooler on this generation of Yankees contenders then perhaps I should have been, but they’ve always seemed to rely on paper thin rotations and relatively fragile offensive stars. They’ve also had strong divisional challenges from the Red Sox and the Rays, although rarely simultaneously. This year, it’s hard to develop a compelling case that they aren’t the best team in the AL on paper as the concentration of top teams seems to have swung to the NL. Their rotation remains dependent on fragile pitchers, but which AL team’s isn’t? The Rays remain the pick for second over the Blue Jays here as I think its easy to underestimate how big the gap between them was last year. The Red Sox contending really would not surprise me, although no one would deserve it less than their entitled fans who would rather ignore Alex Verdugo’s existence than consider that maybe trading a player with one year of control left might make baseball sense.  

AL Central

1. Minnesota

2. Chicago

3. Cleveland

4. Kansas City

5. Detroit

I was all set to pick the White Sox, then I looked more closely at the numbers and concluded that the Twins were a slightly better bet – even before Eloy Jimenez was injured. I have actually consumed more spring training baseball this year due to the circumstances of the times then ever before, and this has caused my opinion of the Indians chances to plummet, which may well be an overreaction. This team will need its pitching to be a strength, and it’s easy to glib and say that the rotation is strong. Upon introspection, it may dawn on you as it did me that they have precisely one starter who has completed an entire MLB season in a rotation. I can’t recall a past Indians bullpen that will rely so heavily on back end arms with good stuff but questionable control, and I actually think Phil Maton may be their best reliever. The offense remains cursed by the franchise’s inability to produce cheap corner bats who can contribute anything. Tribe fans and radio play-by-play announcers alike are contemptuous of the decision to clutch onto Jake Bauers and start him at first rather than put him through waivers, but as uninspiring as Bauers’ past two seasons have been (yes, he didn’t play last year but if you can function as a major league left fielder and were not called up to the horror show that was the 2020 Cleveland outfield, it speaks volumes), Bobby Bradley is not exactly Andrew Vaughn as a first base prospect. A Ben Gamel/Amed Rosario center field platoon? A pair of lumbering Padres castoffs being counted on as key cogs in the offense? And history repeats itself again with the farm system, as through development and trades the Indians have built up a fine collection of middle infield prospects (Andres Gimenez, Gabriel Arias, Owen Miller, Tyler Freeman, Bryan Rocchio) but corner bats remain elusive (a lot rides on Nolan Jones). It’s better than the opposite problem, I suppose, but oddly more frustrating as a fan. The Royals think this is their year because they always think this is their year; I think the gap between them and the Indians is pretty narrow but that doesn’t make you a contender. Will the media en masse ever consider AJ Hinch as a possible feel good story? I would guess not, but what do they know?

AL West

1. Houston

2. Los Angeles (wildcard)

3. Oakland

4. Seattle

5. Texas

Last year, the Astros were both the team hurt most by the sixty-game season and helped most by the expanded playoffs, where they provided evidence that they actually were still a decent team. The starting pitching is scary, the offense is weaker and more fragile than before, but they still stand out in this group of teams. I’ve picked the Angels as a wildcard many times during their upstream swims attempting to get back to Mike Trout’s natural habitat; I’ll probably regret it again, but the ChiSox are no sure bet and the East teams will have a tough schedule to overcome. Perhaps their biggest threat will come from the A’s, still a team that could have a scary rotation (and may have the other kind of scary rotation due to the unreliability of Manaea, Puk, Luzardo, etc.) and a couple offensive stars. This division appears to the epicenter of explicit six-man rotations, with the Angels and the Mariners. Do you think announcers will make a big deal of saying things like: “This is the first time a player has done X in Globe Life Park with fans in the ballpark for a Rangers game?” It seems like a preposterous suggestion but is it really that much more ridiculous than “The Red Sox haven’t won a World Series AT HOME since…”?

NL East

1. New York

2. Atlanta (wildcard)

3. Washington

4. Philadelphia

5. Miami

The 2020 season should have been hard to predict, although for different reasons than 2021. There was no issue with data – the same level of historical statistics was available for 2020 as for prior seasons. The issue of course was that a sixty-game season was subject to a higher degree of variance from expectation than a 162 game season is.

Yet something interesting happened – I did better on my predictions than I ever had before. This was not due to some special insight on my part – I thought the picks I made were pretty obvious and pretty chalky. One of the most interesting things about the 2020 season is how few flukes there were on the team level (of course, this arrogantly assumes that my assumptions were correct – one must acknowledge the possibility that the sixty-game season enabled my poor predictions to appear more accurate than they actually were).

In any event, I was right on five of six division winners, both pennant winners, and the identity of the world champion. I bring this up here because this is the one division I missed on – I picked the Mets, and I’m going to double down.

This division also features the team that I think is mostly to disappoint – certainly would have been, at least, before sabermetric thinking became widely diffused. The Marlins had horrible component statistics last year, should have been bad on paper, look like they are bad on paper again, but made it into the playoffs with a team with a reasonable number of young players, particularly on the mound. It’s exactly the kind of team that it would seem reasonable to think had a breakthrough if you weren’t wise to the fine print.

This is the consensus toughest division, and I don’t disagree – the top four are all real contenders. If you’re a fan of the deserved family of metrics from Baseball Prospectus, bet hard on the Phillies.   

NL Central

1. Milwaukee

2. St. Louis

3. Chicago

4. Cincinnati

5. Pittsburgh

This is the consensus weakest division, and again I concur, although I think the Brewers are very interesting with a high upside collection of pitchers. The Cardinals are getting a lot of buzz for acquiring Nolan Arenado, and I don’t see any reason he wouldn’t bounce back to something resembling his prior form, but in terms of helping them in 2021, I think there were a number of positions where an upgrade would have fit better with the current roster. The Cubs are probably being underrated due to the revulsion to a team that’s been a contender for the past six years seeming to enter a retrenchment, but the offense could still be a force. As shifts have come to the fore, we’ve seen a blurring of the line between second and third basemen, with Milwaukee’s usage of Mike Moustakas as one of the harbingers. Moustakas’ current team is also involved in some interesting infield moves, but bringing back the Howard Johnson as a shortstop strategy is considerably bolder than swapping the Moustakases and Travis Shaws of the world between second and third.

NL West

1. Los Angeles

2. San Diego (wildcard)

3. Arizona

4. San Francisco

5. Colorado

There’s not much to be said about the Dodgers – they are the model franchise of the day, arguably the model franchise of the entire free agency era. It was good to see them finally get a World Series trophy but frankly they deserve more. They would be more than worthy of being the first repeat champions in the last two decades. The Padres are fascinating in their own right, likely doomed to a one-game playoff no matter how much they invest in their roster. One interesting thing about the eight-team playoff structure used in 2020 is that the presumed #1 wildcard team is the only team that would qualify for the playoffs under the old system that clearly has their chances of winning the World Series increase as a result. The division winners all have to play a three-game series to advance under the 2020 system, clearly worse than an automatic berth in the LDS (although if there is a dominant #1 like the Dodgers, the #2 and #3 teams do benefit from a higher likelihood that they get taken out before a potential LCS matchup; it’s not enough to offset having to play a three-game series against a competitive opponent). #5 team would rather be in a one-game playoff with the #4 team than a three-game series, assuming that the team’s regular season records are indicative of their true strength. Of course the #6-#8 teams benefit. Last year San Diego lost the first game of their series with St. Louis; a one-game playoff with Atlanta (as I’m predicting) is a poor reward for all of that investment. The Diamondbacks, Giants, and Rockies are all in that terrible position of being older than you would think (especially San Francisco) and in a division with powerhouses that look to be set up for a few years at least.

WORLD SERIES

Los Angeles (N) over New York (A)

Wednesday, March 17, 2021

Subtweeting Without Twitter, Vol. I

Using a positional adjustment as part of a total value metric (WAR, VORP, etc.) doesn't imply that players can be freely interchanged across positions any more than noting that a pizza and a t-shirt both cost $15 implies an assertion that one can wear a pizza or eat a t-shirt.