When I score a game, I almost always keep a pitch-by-pitch record of the game, unless for some reason I have to juggle watching the game with some other task, and will not have the ability to accurately record each and every pitch. Even when I set out with this as my intention, I often find myself unconsciously scoring the pitches anyway.
My system for tracking pitches only records the basics--whether it is a ball or a strike, and any of the basic subgroups contained within those two categories (intentional balls/pitchouts, swinging strikes, called strikes, fouls). Some people attempt to keep track of pitch locations or pitch types; of course, Pitchf/x has rendered this even more of a chore than it was previously, and some people (hello, nice to meet you) just aren’t good enough at observing pitch locations and distinguishing pitch types to even attempt to put in this level of effort.
The final pitch of a plate appearance is not recorded separately--it is implied by whatever event follows. For example, if a batter draws a walk, I don’t record the fourth ball independently of noting the walk. If a pitch is hit into play, you’ll see a symbol for a base hit, or a groundout, or whatever the case may be. I don’t see any reason to waste another pencil stroke on spelling this out.
The left side of the empty scorebox is used to record balls; the right side is reserved for strikes, and the very top (and on the very rare occasions when necessary, the very bottom), with much smaller letters, is where two-strike fouls are recorded. The order of pitches in indicated by letters of the alphabet--the first pitch is “A”, the second pitch “B”, and so forth.
Balls usually don’t usually need any elaboration--intentional balls/pitchouts are the only common subcategory. I do not distinguish between the two; it is usually pretty obvious which is being employed if you consider the context of the plate appearance and the pitch sequence. An intentional ball of any stripe is simply circled.
The other, much less common alteration needed to balls is the automatic ball, on the rare occasion that the umpire makes that call. The symbol for this is simply a lower case “a” in front of the usual symbol. For example “aD” would indicate an automatic ball called on what would have been the fourth pitch.
There are more modifiers needed for strikes. Called strikes receive no alteration, while a left bracket “[“ is put around the outside of a foul and a left brace “{“ is put around the outside of a swinging strike. The foul symbol is not used with two-strike fouls, since by definition they could be nothing else. Another modifier I use which can be applied to strikes of all kinds (except two-strike fouls) is circling the letter, which is used in case of a bunt attempt.
A couple of examples will hopefully make this pretty clear:

The first pitch (A) is a garden variety ball. The second pitch (B) is a called strike. The pitcher is called for a rare automatic ball (aC) before what would have been the third pitch. The fourth pitch (albeit the third actually delivered) is a foul. The fifth pitch (E) is another foul. The sixth pitch (F) is a ball, and the seventh pitch (G) another foul. Finally, the batter flies to right on the eighth pitch, for which the pitch is not explicitly noted--the occurrence of a flyout is sufficient to demonstrate its existence.

In this plate appearance, the batter shows bunt on the first pitch (A) but takes a strike. The second pitch is a standard ball (B), but the third pitch is a pitchout (C). The batter swings and misses on the fourth pitch (D), and does it again on the fifth pitch for a strikeout.

The batter attempts to bunt the first pitch, but he bunts through it for a swinging strike (A). He attempts to bunt again on the second pitch, but this time he fouls it off (B). He then takes a ball (C), fouls off a pitch (D), and eventually grounds back to the box.
There are several different possible symbols for a strikeout that actually becomes an out in my system--a swinging strikeout, a called strikeout, a strikeout where the putout is something other than catcher unassisted, a strikeout on a missed bunt (that is a swinging strikeout on a bunt), and a strikeout on a two-strike foul bunt. In these examples, I will not include the pitch sequence, since that has already been explained and it would just clutter the scorebox and distract from the out itself.

This is a standard swinging strikeout. The solid dot is my universal symbol for an out; any time the batter-runner is retired at any point, the dot will appear somewhere within his scorebox. This makes it much easier to quickly see how many outs there are in the inning, and also eliminates some potential confusion in cases in which a certain code could indicate an out or could indicate something else.

And the inscrutable backwards K for a called third strike.

Sometimes the scoring on a strikeout is something other than the standard catcher unassisted. By far the most common is the catcher throwing to first for the out (23), although there are other possible and weird ways for this to occur.

This is the symbol I use for a foul bunt with a two-strike count, resulting in a strikeout. As I’ll show later, the squiggly line is my symbol for a bunt on a ball in play, so the symbol applied to that of the strikeout has a clear meaning.

If a bunt is attempted but missed for a third strike (that is, the batter offered but did not make any contact), then the brace that indicates a swing for a non-third strike is included in the symbol above to distinguish it from the more common third strike bunted foul.
Tuesday, May 03, 2011
Scoring Self-Indulgence, pt. 2: Scoring Pitches and Strikeouts
Thursday, April 21, 2011
Wayne Winston's Mathletics
The "book reviews" on this blog are almost always a day late and a dollar short. They are written and published long after the book, and my comments about them usually don't amount to a review but rather as a springboard from which to discuss other topics. This one is no different.
Wayne Winston is a professor of Decision Sciences at Indiana University's business school and a former consultant to the NBA's Dallas Mavericks. He published Mathletics in 2009 with the tagline "How Gamblers, Managers, and Sports Enthusiasts Use Mathematics in Baseball, Basketball, and Football."
If you are a regular reader of this blog or similar material, do not buy this book expecting to learn a lot of new things about sabermetrics. The sabermetric material is fairly standard, rudimentary type material--introductory-level discussion of run estimators, park factors, replacement level, the base/out table, win expectancy, and the like. I would also not recommend it to a novice, not because it is poor (there are elements I like and dislike, as I'll discuss below), but because there are better resources out there--internet primers, Bennett and Fluck's Curve Ball, and Lee Panas' Beyond Batting Average among others.
I am not particularly well-read on either football or basketball quantitative analysis, so I cannot definitively state the level of Winston's discussion on those topics. My guess is that the football discussion is fairly basic (with the caveat that football analysis as a field lags behind apbrmetrics), but that the basketball material is much stronger. It is certainly obvious from the writing that basketball is Winston's passion, and that the adjusted plus/minus ratings are a particular favorite.
Winston's writing is not particularly strong--he writes like someone whose favorite class was math (as do I). There are some minor slip-ups in the baseball discussion; these won't mislead the reader, but they also reflect the pedestrian nature of the material:
* Winston includes a formula for estimating batting outs that accounts for ROE by putting a multiplier on at bats. But this applies the adjustment to all at bats, including those in which we know a batter did not reach on an error (hits) and those in which the likelihood was very small (strikeouts).
* He refers to Keith Woolner's statistic as VORPP--Value Over Replacement Player Points. This makes sense in that he applies the replacement level concept to WPA points, but he also refers to Woolner's run based version as VORPP. Additionally, he credits the concept of replacement level to Woolner. In reality, Woolner did much to popularize replacement level, but the concept did not originate with him.
* Similarly, he credits the concept of park factors to Bill James. James had much to do with popularizing the notion that statistics could be corrected for park effect, but if any single person is to be credited with the concept, Pete Palmer would be an easy choice.
* There is a chapter that discusses player improvement over time by comparing annual performance, but it does so without even really addressing aging and survivor bias.
* The discussion of strategy is fairly bare-bones and deals only with basic estimates based on a standard run expectancy table.
There are positive things of similar magnitude to the list of negatives--for example, while he uses Runs Created, he explains that a theoretical team construct is necessary to make accurate player comparisons. As a whole, the baseball portion of the book is adequate without being excellent for a novice and a yawn for those well-versed in sabermetrics.
Being a novice myself when it comes to football and basketball analysis, I found the discussion in those chapters much more interesting. Focusing on a couple interesting football tidbits, Winston offers a version of the famed two-point conversion chart that incorporates the expected number of possessions remaining in the game. There is also a formula for the probability of a successful field goal in the NFL based on distance that I found interesting, although the model produces results that are clearly too high for very long kicks.
There is also a discussion of quarterback ratings, which have always interested me. Like every other sane person, Winston has little use for the NFL system, focusing his discussion on Berri's rating from Wages of Wins and his own adaptation of Brian Burke's regression of team categories against team wins. Isolating the categories from Burke's equation that can be related directly to individual quarterbacks, Winston offers the following as a quarterback rating:
1.543*(Yards - Sack Yards)/(Attempts + Sacks) - 50.0957*(Interceptions/Attempts)
If you factor out and ignore the 1.543 coefficient, and change the second quantity's denominator to (Attempts + Sacks), this can be rewritten as:
(Yards - Sack Yards - 32.47*Interceptions)/(Attempts + Sacks)
In this form, Winston's rating is very similar to a number of rating formulas, including the NEWS rating published by Bob Carroll, John Thorn, and Pete Palmer in The Hidden Game of Football:
NEWS = (Yards - Sack Yards - 45*Interceptions + 10*Touchdowns)/(Attempts + Sacks)
Breaking into editorial mode and stepping away from Mathletics for a moment, the treatment of a touchdown pass can be thought of as somewhat analogous to the sacrifice fly in baseball. The comparison is strained as touchdown pass is always a positive play from any perspective, while a sacrifice fly might actually reduce run expectancy.
A fairly large number of touchdown passes occur on short passes. Suppose a quarterback completes a three-yard touchdown pass. This will actually reduce his rating in Winston's ranking, as the quarterback's rating prior to the touchdown will be higher than three. By giving a positive weight to all passing touchdowns, one could ensure that a touchdown pass always increases ranking.
However, in doing so, one gives special treatment to the touchdown because it is a tracked category (like sacrifice flies). However, one could also track "sacrifice grounders" or "first down completions". These theoretical categories would also be cases in which a positive or somewhat positive outcome was achieved, but the statistics treat it as a negative (a batting out or a reduction of the passer's rating, assuming the completion was short). Giving special treatment to the recorded categories can thus be seen as unhelpful and biased by particular types of players that might be predisposed to one or the other.
Moving back to the book, most of my comments to this point have focused on the negatives. However, there are three things that Winston does really well:
1. Winston provides downloadable spreadsheets for many of the examples. This allows the reader to follow along with the work and to learn how to carry it out in Excel. Many of the Excel steps are explained in the text as well.
The drawback to this is that some of the why behind the math is glossed over in favor of a quick Excel solution. Winston's rating system for NBA and NFL teams basically boil down to finding the best-fitting solution for a system of linear equations to predict the point margin in each game. Winston doesn't explain the math in that manner, though, instead just explaining that the Excel solver is used to minimize error. While this gives the reader enough detail to produce their own ratings, and no one is actually going to solve hundreds of equations, I personally prefer a stronger emphasis on the underlying math.
2. The bibliography is excellent, as it includes not just a list of sources but descriptions of what they offer. For example, this is the description of Phil Birnbaum's Sabermetric Research blog:
This is perhaps the best mathletics blog on the Internet. Sabermetrician Phil Birnbaum gives his cogent review and analysis of the latest mathletics research in hockey, baseball, football, and basketball. This is a must-read that often gives you clear and accurate summaries of complex and long research papers.
3. Winston's description of Birnbaum's blog provides a nice transition into discussing the best thing about his approach. While Winston has excellent academic credentials (he is a professor of Decisions Sciences at Indiana and earned a PhD at Yale in Operations Research), but he does not beat you over the head with it. In fact, I don't think that his doctorate is ever explicitly referenced.
In any event, Winston mixes the research of other academics into his text, but he gives plenty of space to amateurs as well. Some academics that enter the sports arena seem to thumb their nose down at anyone who doesn't hold an advanced degree or a teaching position. Winston is not one of them. He even used one of Birnbaum's posts to offer a counterpoint to an academic paper on the NFL draft.
Winston's book provides a great example of how sabermetric knowledge generated by academics, amateurs, and everyone in between can be integrated, and how all parties can respect and learn from each other. It also gives analysts specializing in each sport a window into the work being done on other sports. Thanks to those attributes, Mathletics is a worthwhile read.
Wednesday, April 13, 2011
Comments on Baseball Prospectus 2011
At some point it becomes bad sport to write the same thing about an annual book--if there’s a certain characteristic of the book that you find yourself dissatisfied with several years running, it might be a you problem. It’s one thing to decide that a certain book is not for you; it’s another to continue to believe that it will when it’s obvious that the writers have something else in mind.
Much of what I could say about the Baseball Prospectus annual for 2011 is the same as I said about in 2010, and 2009…and so I’ll try to avoid saying it again. By now, it’s clear that BP is what it is, and that can either be a great thing or a bad thing or a mostly good thing, depending on your perspective. My perspective is that it’s mostly a good thing--the redeeming qualities of the book outweigh its flaws fairly easily from my perspective.
I still felt compelled to jot down a few comments on the book this year because I might have been a little unfair in nitpicking a few things in the past. Now that there is a lot of new blood on board, it’s more apparent that some of the issues (like stats not matching up between the comments and the data directly above) are systematic, and probably endemic to producing a book of this kind. To put together a tome of that size in a few months is a massive undertaking, and there are thousands of moving parts, so expecting them all to be dialed in to the same setting is unrealistic.
The cover still has the infamous phrase that I will not repeat about PECOTA; this is obviously out of the hands of the writers. They do redeem the cover with a great caption under the little photo of Albert Pujols.
That being said, I do have one major bone to pick with the new, slimmed down statistical offerings. It’s great that they stopped doubling up on metrics that measure the same thing (in the past, there have been simultaneous displays of VORP and WARP, or EqA and MLVr), and with one glaring exception the new stat lines still manage to give you most of the key metrics. That glaring exception is the lack of any kind of component ERA (or RA, which I’d prefer anyway) figure for pitchers.
It’s not simply a matter of limiting your choice to vanilla, while having to leave chocolate, strawberry, and cookies and cream aside (after all, there are a lot of flavors of component ERA). There is none whatsoever. Instead, BP has listed Fair RA, which is a fine metric constructed by Colin Wyers and the primary input for pitcher WARP. But if the choice is between having Fair RA and a component ERA in a book that is largely aimed toward predicting performance in 2011, it’s not a choice at all. Sticking with metrics under the BP umbrella, peripheral ERA and SIERA would fit the bill.
Of course, if I could strike any category from the pitcher stat line to clear space, it wouldn’t be Fair RA--it would be W-L or saves or WHIP. But since a big target audience for the book is fantasy players, that is not an option. However, it leaves everyone (including fantasy players) without a backwards looking metric that gives us the best estimation of how the pitcher’s overall effectiveness in the past. I certainly hope that they will figure out a way to include Peripheral ERA or SIERA or something similar in the 2012 edition.
PECOTA is in good hands with Colin Wyers, and I’m sure there are still some bugs to be worked out, so please take this comment as more amusement than criticism: some of the PECOTA comps seem way off. I’m sure this happened in the past, and I didn’t bother to make note of it, but two players that really stood out to me were Gregor Blanco and Nick Franklin. Blanco’s top comps are Richie Ashburn, Kenny Lofton and Freddy Guzman. One of these things is not like the other, and two of them are nothing like Gregor Blanco (Lofton was still in the process of breaking out, but had already established himself as clearly better). The Franklin comps are more understandable since he’s a younger player with less of a track record, but it’s still an odd juxtaposition to see a player ranked as the #44 prospect in MLB while his top comps are identified as Adrian Beltre (ok), Hank Aaron and Willie Mays.
There are only a few team entries that have extensive sabermetric (as opposed to applied sabermetric) content. One of these is the Arizona entry, and sadly I have a bone to pick with it. The author accepts the mainstream view that Arizona’s copious strikeout totals in recent campaigns had doomed their offense. He (or she; I still maintain it would be more interesting to know which author is responsible for the team entry) asserts that “when the majority of the lineup falls prey to empty at-bats of this sort, highly volatile run-scoring can result.”
While there have been some studies done on the relationship between shape of offense and scoring distribution, I am personally unaware of any comprehensive or well-established enough to make a statement like that without the need for supporting evidence. The only statistic brought in to support that position is that Arizona scored three or more runs per inning as much as the NL average, but scored two or less more often.
That is a very odd and not particularly helpful way to break down innings, because it lumps scoreless innings in with one and two run innings. To be absurd for a moment, if an offense never scored three or more runs an inning, and scored 0-2 in 100% of their innings, but 40% of those were one run and 10% were two runs, they would average a healthy 5.4 runs per game. It is true that Arizona scored in a smaller proportion of their innings than did the average NL offense--25.9% of Arizona innings resulted in a run scored compared to 26.5% for the league as a whole. But Arizona was more likely to have a multi-run inning (12.4%) than the average NL team (12.2%).
Another odd thing about this perspective is that it makes the inning the unit by which scoring volatility is measured. It’s true that the best perspective from which to understand how runs are scored is the inning level, since the events that transpire in each inning is independent of those that occurred in previous innings in terms of scoring in runs (I hope it’s clear that I’m talking about baserunners and outs from one inning affecting each other, not lineups turning over and pitchers being removed and the like, but you never can tell) but from a win/loss perspective, it is the run distribution per game that is crucial. Admittedly, the two are very closely related, but any time you extend the time period over which such volatility is projected, its impact is reduced.
One crude but simple and reasonably sensible way to consider the win value of a team’s per game scoring distribution is a method that I call Game Offensive Winning Percentage (gOW%) and have published here for the last three years. It is based on a Bill James idea; instead of estimating an OW% from average runs scored per game, use the team’s actual distribution of runs scored. If in a given season teams that score one run win 11.8% of the time (as they did in 2010), then credit the offense with .118 wins for each game in which they score exactly one run. Repeat for all scoring levels and average and you have an alternative OW%.
There are of course flaws with this method--the unit of games doesn’t always represent the same things (i.e. there are not always 27 outs per game), the use of the actual W% by runs scored in any given season is subject to sample size fluctuations, there is no adjustment for park, etc.--yet it’s still reasonable to think that if a team’s run distribution was particularly unusual, it would manifest itself in a comparison of gOW% to standard OW% based on average runs per game (in this case, without a park adjustment so as to better match gOW%).
The Diamondbacks led the NL in strikeouts in 2009 and 2010 and were second in 2008. In 2007, they ranked eleventh (and made the playoffs, see!), so those three seasons are the relevant high strikeout seasons for the team. In 2008, Arizona’s gOW% was .485 while their OW% was .479--considering their run distribution rather than just their average suggests an additional win. In 2009, it was .484/.483--no difference. In 2010, the split was .492/.502, which is -1.6 wins. So for the three years considered together, the net total is -.5 wins.
Of course, this does not conclusively demonstrate that Arizona’s offense was as efficient as a typical offense with their scoring average, and it certainly doesn’t allow us to make any statements about the effect of high strikeout offenses generally. However, neither does anything offered or referenced in the BP essay, yet the author chose to make much stronger assertions than I would dare to here.
My comments on strikeouts should not be taken as a negative judgment of the book as a whole--my book “reviews”, such as they are, generally serve as an opportunity to discuss issues raised by the author rather than to offer a summary judgment on the book itself. By now, you already know whether BP is a book for you or not.
Tuesday, April 05, 2011
Scoring Self-Indulgence, pt. 1
When I have occasion to write something on paper, I usually use a pen. It’s easier that way--ball-point pens are ubiquitous and cheap; you can sign things with them; and now that the hideous scourge of blue ink has faded a bit, they no longer result in an assault on one’s sensibilities every time they are used (okay, that last one should say “my sensibilities”). In truth, I like pencil better, specifically a mechanical pencil with .5 lead. I use the real cheap Bic ones exclusively, and have for years--you know, the ones that are supposed to be disposable, but you can hold the clicker down and push the replacement lead in through the top. You can get ten of them at Wal-Mart for $2.
The pluses of the ball-point pen allow me to save that favorite writing utensil for only the most important tasks, ones that just can’t be entrusted to the terrifying permanence of ink. For most of the winter, it sits undisturbed on my bookshelf or in a pencil holder or wherever--but sometime in March, I have occasion to take it out and put it to use, and I don’t stop until mid-autumn.
You have probably surmised by now that the important task to which I refer is scorekeeping. Yes, the existence of internet gametrackers have made the collection of data for one’s own perusal something less than a necessity if one would like access to real-time information on a game, and to the extent that people do want to keep their own score, electronic applications are pushing pencil and paper aside. And admittedly, those of us who keep score not just at the ballpark but in the privacy of our own homes have always been a rare breed and prime targets for the nerd label.
Still, I have no intention of giving up scorekeeping in the foreseeable future. It is still true that if you want something done right, you have to do it yourself. GameDay may have all of the information I need, but it (cannot yet at least) be customized to display it in the exact manner I have become accustomed to. If you want to save it for posterity, a GameDay printout lacks any sort of sentimentality whatsoever. And I might be part of a dying breed, but if I want to give my full undivided attention to the ballgame, the last thing I need to be doing is puttering around on the computer between pitches.
If this reads as a half-hearted defense of scorekeeping, I have accomplished what I set out to do with this post. For one thing, I don’t really need to justify my hobby to you; I just feel compelled to put in a good word for the practice every once in a while. I’ve never understood why announcers sometimes feel compelled to give you basic information about the sequence of plays in a game--information that they are tasked with providing--by prefacing it with “If you are keeping score at home…” Of course, this is a hanging curveball set up for the announcing partner, who gets to jump in and make a snide comment about what kind of deviants would be doing that. Considering that those of us that keep score are the least likely subset of fans to turn the game off when it’s 14-2 in the bottom the eighth...
But the other reason that it can be difficult to espouse the virtues of scorekeeping is that scorekeeping is a very personal pursuit. Everyone has their own technique, their own special symbols built around the familiar position numbers that have united the vast majority of scorecards from the 1890s or so on. (Except for the early twentieth century occasional flip-flopping of 5 and 6 for third base and shortstop). This makes it difficult to generalize--I might say that I love keeping score because I could quickly get a precise count of how many balls the hapless Ranger pitches had thrown while Neftali Feliz waited in the bullpen…but your scoresheet might not tell you that. Instead, it might tell you who won the sausage race.
There are displays of the variety and innovation in individual scorekeeping out there online, but not to an extent that I consider sufficient, so last year I asked people to send me their scoresheets for posting on my scorekeeping blog, Weekly Scoresheet. Several people graciously accepted my invitation, but I was foolish enough to make the initial request during the offseason, when even compulsive scorekeepers weren’t particularly likely to have an example sitting around. (*) So if you’re interested in sharing now, please send me an email.
Weekly Scoresheet has a whopping total of six subscribers on Google Reader, which is completely understandable--a personal scorekeeping blog is a vanity blog, plain and simple. Unfortunately, I haven’t been updating it recently because I no longer have a scanner at home, and it just isn’t a big enough priority for me to buy one even though this is 2011 and they are cheap. Eventually, Weekly Scoresheet will be back in full swing to bore the five people who read it with my own chicken-scratched records of ballgames.
In the meantime, though, I’ll be using this space to occasionally run a tutorial on my scoring system and walking through a sample game from the 2010 season. Calling it “my scoring system” is a misnomer--there's certainly nothing groundbreaking about it and most of the symbols are drawn from other people’s systems--but the great thing about scorekeeping is that the precise combination of data you record, codes you use, and the like is fairly unique. I would not encourage anyone to learn to score a game in the way I do, not because I don’t think it’s a decent system but because I would encourage you to organically develop your own that fits your needs and interests as a baseball fan. As such, an explanation and tutorial is ultimately just a way to fill up space and pad the post count. I’ll enjoy writing it; if you enjoy reading it, then much the better.
(*) Sadly, I have to admit that I spend a not insignificant free time in February doodling scoring of imaginary games, using Excel to design some new scoresheets that I’ll never use because I continue to use the same basic sheet (and I do mean basic) that I have for over a decade, wondering why the calendar can’t turn to March so that spring training games can be scored, and other such pursuits.
I will begin simple, with a look at one of my blank scoresheets--it's just a bunch of empty blocks in a 9x9 grid. I originally made this in the DOS Brief text editor prior to the 1998 season, using a lot of “_” and “|” symbols. The lines weren’t solid, so I eventually traced them down for 1999 or so. Later I would scan it as a PDF and touch it up a little bit in Photoshop, but it still has non-perfect lines which to me grants it a little character that you don’t get from using a computer to draw the lines precisely. I have a couple facsimile versions created in Excel (which I use now for any new scoresheets I make--it might not be as capable graphically as some other programs but making grids is something that spreadsheet software does very well), but there’s nothing like the original article for me.

* Why a 9x9 grid? You do realize that the average team doesn’t even use half of those scoreboxes in a game, right?
One of the great things about the Project Scoresheet is that it pioneered the use of numbered boxes rather than a box for every batter to hit in every inning. This was a great way to conserve space, but it also makes it harder to view the inning as a standalone unit. I prefer to see each inning on its own. Yes, a team batting around is a mild inconvenience, and splits up an inning into two columns, but I’ve never understood why some people freak out about and start crossing off the inning headings and pushing each inning down a column.
* No room for statistical lines (AB-R-H-BI and the like)?
Nope. I don’t think that standard compiled statistics for a single game tell you much of anything, and the sort of boilerplate box score is something that is very easy to obtain online (although some of them aren’t so accurate). I’m already sacrificing space by using the 9x9 format; I don’t want to waste any more on stat lines. Plus, if you fill them in as the game goes on, it’s another distraction and you have a bunch of ugly tallies on your sheet. If you wait until the game is over to finish your scoresheet, then it’s just work.
* No diamonds?
Nope, I’ve never liked them. They’re great if you just want a quick snapshot of where the runners are, but if you’re trying to record a lot of detail on how runners advanced, they get in the way. I split the box up into four corners for each of the bases, but I don’t think a visual aid like a diamond is necessary to accomplish this.
Wednesday, March 30, 2011
2011 Predictions
There are many great things to be said about internet publication. It’s free, it’s instantaneous and any boring person like myself can get his thoughts out there and have them found by interested parties. The unfortunate thing about it is that any idiot can find it too. If you write in print or for a price, idiots still read your work (or worse yet, get secondhand accounts that substitute for reading), but there is a higher barrier to entry than a lucky Google search.
No matter how many disclaimers you put on a post, the reader can always ignore them. The reader can quote you out of context by means of a simple copy and paste that any semi-coordinated five-year old could manage. No matter how careful you are, there is a decent chance that your disclaimers will be ignored and your words will be turned on their head.
Making the disclaimer more strenuous will do no good, naturally, but it’s the only option available to the blogger. If you find this site via a Google search and fail to read and understand the next few sentences, that's on you, not me.
These predictions are offered in the spirit of fun, not science or anything resembling science. They are one man’s opinion, one man’s wild guess at an unknowable future. If you do not believe that future is unknowable, if you think that you are blessed with some special insight that allows you to predict the outcome of pennant races, you are likely incapable of understanding this point, but nevertheless: even if you could predict exactly how many runs a team would score and allow over the course of the upcoming season, you still could do no better than predict their win total to within a standard error of four games.
Of course, some folks predictions will be more accurate for a single season. Some will be better over the long haul, too, although expecting a persistent and consistent performance is folly. It is quite possible that I am poorer than most at this exercise.
Understand, though, that the picks that follow are the product of one man’s feeble mind. They do not in any way, shape or form reflect the predictions of “sabermetrics”. They are not based on any sabermetrically rigorous procedure; I do some very crude estimates of team runs scored and allowed based on freely available projections, but then I substitute my own opinion wherever I feel. If you want predictions based on a more rigorous sabermetric procedure, you need to visit Baseball Prospectus or the Replacement Level Yankee Weblog or somewhere, anywhere else--not this blog.
Another point I make every year but which is hopelessly lost on the mental midgets who invariably stumble upon this page is that the format distorts my true feelings. I have limited myself to the format of ordinal standings rank not because it is a better format, but because to do what is more telling, offering expected wins and probabilities, requires the use of a rigorous procedure to be remotely credible, and then it becomes harder to shrug it off as a fun, seat-of-the pants exercise.
Given the constraints of the ordinal prediction format, one is forced to predict which team will win the NL Central, and which will finish second, and all down the line. In reality, I have no idea who will win the NL Central. I believe that three clubs--Cincinnati, Milwaukee, and St. Louis--are essentially equal in their chances. A fourth, Chicago, could certainly win if things broke their ways. And while it may be unlikely, one cannot give true zero odds to the possibility of a miracle in Houston and Pittsburgh. Add it all up, and I wouldn’t be willing to give you better than a 35% chance that any single team wins. But making an ordinal prediction gives the impression that the predictor believes that the team listed first has an expected outcome of first place. The intelligent reader and prognosticator understand the difference and don’t need to have it spelled out.
Anyway, I’ve done all I could, spilling 650 words on why what follows should not be taken seriously. Someone with Google will invariably stumble upon this and be unwilling or unable to understand anyway. Oh well:
AL EAST
1. Boston
2. New York (wildcard)
3. Tampa Bay
4. Baltimore
5. Toronto
I agree with the consensus view--the Red Sox look to be the best team in baseball. Their offense should be great, their starting pitching might have only one truly safe bet in Jon Lester but also features four guys with the potential to pitch really well, and their bullpen has three big arms (albeit perhaps with two empty heads) at the back. Too much is being made of the Yankees' pitching concerns, I believe. There are a lot of teams expected to contend with question marks at the back of their rotation, but since we're conditioned to see New York run out four guys with track records, it looks worse than it is. I was very tempted to pick the Rays second, but it would be a dishonest pick for shock value. This should still be a very competitive ballclub, and there a couple divisions I probably would pick them to win. The mainstream is spending too much time fretting about losing Pena, Garza and Bartlett and not enough recognizing that their performances in 2010 were far from irreplaceable. Crawford is obviously a much bigger loss, as is their bullpen's splendid 2010 performance which couldn't be replicated even if they were all still in St. Pete. The Orioles are not nearly as close as they seem to think they are, but I thought they'd be better last year and meeting those expectations could be enough for fourth place. I like Alex Anthopoulos' moves as much as the next guy, but I liked much of what Jack Zduriencik did last winter too. I don't think the Blue Jays project to be a better team than they were last year, and their homer-happy offense is almost surely unsustainable. If you had to fall into the silly cliché trap and pick “this year's Mariners”, you could do worse than Toronto.
AL CENTRAL
1. Chicago
2. Minnesota
3. Detroit
4. Cleveland
5. Kansas City
I've had it out for the White Sox organization for a long time, mostly focused on their manager. It's more a personality thing than a baseball thing, but I'd be lying if I said that I've always been able to successfully separate the two. So it is not with any sort of relish that I predict they'll win the AL Central. But their starting pitching is excellent, and while I wouldn't want their 2013 offense, their 2011 offense should be fine. The Twins did nothing, which is usually not a great sign. If Justin Morneau is healthy, I'm not sure he's a great bet to provide much more value than he did in 2010, when he was arguably the best hitter in the league for half a season. My crude spreadsheet actually gives Minnesota an insignificant edge over Chicago, so I see these two clubs as quite close. The Tigers are spending money without any particular target in mind, or at least that is the impression they give off. They are a contender in this division, but they represent the second tier. The second tier is small because it does not include the Indians or the Royals. Neither pose much of a threat, and I won't rehash my thoughts on the Tribe here. Dayton Moore's farm system is a universally-acknowledged jewel, but the man has yet to provide much evidence that he is capable of constructing a winning major league team, although this wasn't really the year to try, One can easily imagine five-star prospects surrounded with the Jeff Francoeurs and Melky Cabreras and Pedro Felizs of 2014 (that would be a fun exercise--who is the next Jeff Francoeur? Who was the last Jeff Francoeur?)
AL WEST
1. Texas
2. Oakland
3. Los Angeles
4. Seattle
The Rangers are my reluctant pick--I love to pick against pennant winners without sterling regular season records, because the mainstream halo effect for playoff success is far too strong. Like the talk of the Twins dangling Liriano, the decision to keep Feliz in the bullpen gives off a sense of complacency and overconfidence. However, Texas does look like the strongest club on paper. Of course, games aren't played on paper, and so if the great clubhouse influence of the sainted team-player Michael Young is dispatched, they will plummet to 100 losses. The A's made an effort, but their offense still makes them hard to pick. They also don't have a strong Buster Posey candidate to emulate the 2010 trick of their baymates. The big splash Tony Reagins promised for the Angels was fulfilled if you think about “big splash” in terms of Olympic diving. The Pythagorean fairy better bring some extra pixie dust. The Mariners have to score more runs, don't they? Don't they?
NL EAST
1. Atlanta
2. Philadelphia (wildcard)
3. New York
4. Florida
5. Washington
Four aces = unbeatable! The pitching fetish that still looms large in the conventional narrative demands that tribute be paid to the Phillies, but would it be sacrilege to point out that pitchers often get hurt, and an offense built on a bunch of players on the wrong side of thirty whose best player is a major injury question might be a little vulnerable? Just checking. The Braves have been my irrational NL pick for some time now; since they finally rewarded my faith with a playoff berth, there's no reason to stop now. The bullpen is due for some regression, but the starting pitching looks fine and they should score some runs. The Mets look like a .500 team, which means the ratio of lamenting how bad they are to reality will be way out of whack. The Marlins seem stuck in neutral, even if they hadn't traded Uggla; even if you believe in success cycle theory, it's no one's birthright to win the World Series on a six-year cycle. This might be the MLB division with the most delusional owners; the Mets apparently put all their eggs in one basket, Loria thinks his world title is two years overdue, and the Nationals think that Jayson Werth + Stephen Strasburgh + Bryce Harper = 2012 contention without a lot of downside risk. Good luck to all.
NL CENTRAL
1. St. Louis
2. Milwaukee
3. Chicago
4. Cincinnati
5. Pittsburgh
6. Houston
My crude projections have some kind of borderline irrational love for the Cardinals. Perhaps it's the fact that they ignore fielding; perhaps there's some truth there. Picking them to win after the Wainwright injury in an already close division may seem foolish, but what the heck? At least I can point to the fact that it's not alone—PECOTA also puts St. Louis on top, albeit by an insignificant margin. The Brewers have shown a tendency to go for it with gusto by trading for starting pitchers, and while I wouldn't want my team doing the same, it could very well pay off in this division; and for the moment, fate has smiled upon them. The Cubs probably aren't as close to winning this division as the Garza trade suggests they think, but a lot went wrong for them last year (some of it of their own volition, mind you), obscuring the fact that in 2009 they were kind of in the mix. My own intuition tells me that the picks for this division are scrambled, but the spreadsheet really doesn't like the Reds that much, and I can see why. Their offense still has issues at short and in left, Scott Rolen was a big contributor in 2010 but at this stage isn't the most reliable guy around, Joey Votto isn't really Albert Pujols Jr., and I'm not really a Drew Stubbs believer. They don't really have a leadoff hitter (which I point out not because the leadoff role is particularly important but because they either don't have or don't trust high OBA guys without power), and you still have to wonder about Dusty's ability to push the right the buttons if things don't go according to plan. The starting pitching has solid depth, but they also lack front line starters barring a Cueto or Volquez breakout. They are certainly the team I've picked fourth in a division that I think has the best chance to win it--again, that's the inherent peril in using an ordinal standing prediction approach. It should go without saying that they're much closer to St. Louis on paper than they are to Pittsburgh. The Pirates should not lose 100 games again, and the Astros might be as good of a bet as any team in the game to do so.
NL WEST
1. San Francisco
2. Los Angeles
3. Colorado
4. San Diego
5. Arizona
I really don't like picking the Giants. Winning the World Series doesn't wipe the fact that their playoff berth was not secured until the final Sunday of the season out of existence, and their offense is still quite suspect. However, the Rockies didn't do much to improve; I wouldn't be surprised to see serious regression from Carlos Gonzalez, and my crude projection attempt sees them as a .500 team. The Dodgers could surprise some people since expectations are not particularly high. Their pitching staff should be one of the best of the NL (at least before park adjustments), but there are a number of weak/questionable spots on offense. The Padres obviously traded away their best player without much coming back to help them in 2011; still, they should be respectable. The Diamondbacks may think that shedding strikeouts is the magic elixir to increased offense, but you still should expect to score more runs with Mark Reynolds as your third baseman than Melvin Mora.
WORLD SERIES
Boston over Atlanta
While I think they are far from a juggernaut (just as I thought during last year’s postseason), I would have to admit that the Phillies are the team most likely to win the NL pennant and the NL East. However, my estimation of their chance to do so is so much lower than that of the mainstream that to pick them to win and thus replicate everybody's prediction would be boring. I almost feel the same way about Boston--the scattered talk about winning 110 games is irrational. However, I think Boston stands out from the pack more than Philly, and picking against your head twice is just a waste of everyone's time. My crude standings projections have the Giants second in the NL behind the Phillies, but picking them would be even worse in terms of overvaluing the previous postseason and assuming repeats, so the mantle falls to the Braves.
AL Rookie of the Year: SP Jeremy Hellickson, TB
It doesn't feel like he should be eligible, but he is.
AL Cy Young: Jon Lester, BOS
Let's try this again. I picked him last year and he had a fine season, but not a spectacular one.
AL MVP: 1B Adrian Gonzalez, BOS
I always like to pick a newcomer for MVP if possible (voters love shiny things), and Fenway should help the mainstream perception of his performance if nothing else.
NL Rookie of the Year: 1B Freddie Freeman, ATL
Having a job is half the battle.
NL Cy Young: Clayton Kershaw, LA
I was going to pick Zack Greinke, but missing a few starts is a hurdle to overcome.
NL MVP: SS Troy Tulowitzki, COL
You know that the writers are just dying for an excuse to give this to Ryan Howard.
First manager fired: Jim Riggleman, WAS
Spending all that money on Jayson Werth leads me to conclude that the powers that be in Washington think they're a stronger club than they are.
Best pennant race: NL Central
Worst pennant race: AL West
Worst team in each league: KC, HOU
Most likely to go .500 in each league: OAK, CHN
Team in each league most likely to disappoint mainstream consensus: DET, CIN
Team in each league most likely to surprise mainstream consensus: TB, LA
Most obnoxious stories of the year: Michael Young, Pujols' contract status, Nolan Ryan taught the Rangers pitchers to do X, Josh Hamilton, whether various starters turned relievers should be used in a sane manner (Chapman, Joba, Feliz)...basically the Texas Rangers, win or lose.
Tuesday, March 22, 2011
Historical Park Factors, 1901-2008
I have posted an updated spreadsheet with park factors for all teams, 1901-2008, as a Google Spreadsheet. These are five-year park factors, calculated in the same manner I describe on this page.
The guiding philosophy was to try to include as much data as possible. If there are five possible years of data to be used for a park, they will all be used, even if four of the seasons were in the past or in the future. The source of the raw data was KJOK’s excellent park database for past seasons and various sources (most notably Baseball-Reference) for recent seasons.
I treat a park as new if there are major changes to the dimensions, but I did not by any means do a complete historical survey to find out when those changes have taken place, so some that probably should have been treated differently are not. If you have specific data on when a change should have (or shouldn’t have) been made, feel free to leave a comment and I will try to incorporate these changes when I update the chart some time in the future.
Additionally, when a team moves, and a new team immediately moves in (for example, the Senators of ’60 and ’61), this is treated as a new team. Also, in cases in which teams have played a significant (which I defined as around ten or more) number of games in a different stadium in the same year, those years are treated as being a new park (an example is the Dodgers playing games in New Jersey the two years before they moved from Brooklyn). Whenever a “new park” of this sort is established, when the old order is restarted it is treated as another new park.
The reason the park factors are only shown through 2008 is that my ideal data set is two previous years, the year in question, and two future years. For most of the parks active in 2009, we will after 2011 be able to fill this dataset, and so I don’t want to publish a park factor now and change it later. However, there are parks where the 2008 or 2007 factors are not yet settled because they are new and there are not yet five years of data available. In these cases, I have listed a PF but marked it as one that will change in the future (this is indicated with an orange shading; park factors for the first year after a switch are in pink text).
Now I will give an example of how I chose the years to be considered in figuring the PF. Suppose we look at the Diamondbacks, who have played in Bank One Ballpark since 1998. In 1998, we have no previous data, but there is four future years of data, so the sample is 1998-2002. For 1999, there is one previous year, so we also look at three future years, and get 1998-2002. For 2000, there are two previous years, so we use two future years, and have a sample of 1998-2002. This is now in the ideal format--the year in question, plus the two immediately prior and future years. Of course, in 2001, we use the two previous years (1999 and 2000), and two future years (2002 and 2003), making the total sample 1999-2003, and it will continue in that manner until something changes.
Let’s also consider the end of the Braves’ tenure in Fulton-County Stadium. The last season there was 1996. For 1994, we have two previous years (‘92 and ‘93) as well as two future years (‘95 and ‘96), so we use 1992-1996. For 1995, we have just one future year, so we use three previous years, and also use 1992-1996, and the same for 1996.
In the previous iteration of these park factors, there were three recent parks for which I needlessly inserted a changepoint and thus changed the factors for the surrounding seasons. These teams were pointed out by Terpsfan101, and I have corrected their PFs in this edition. They are Detroit, 1994-1999; Minnesota, 1989-1995; and Seattle, 1993-1996.
Tuesday, March 08, 2011
2011 Indians “Preview”
If you close your eyes and dream just a little bit, you can see the Indians contending in a division without a clear favorite. A healthy Grady Sizemore and a healthy Carlos Santana teaming up with the overlooked Shin-Soo Choo to form a secondary average all-star wrecking crew at the heart of the offense. Maybe Matt LaPorta lives up to his potential, just a little, and provides average first base output. Michael Brantley is able to get on base, and Asdrubal Cabrera fields well, and second and third base are not complete black holes. Meanwhile, the bullpen is anchored by Chris Perez, with Rafael Perez and Tony Sipp combining to give the Tribe two solid left-handed relievers. A couple of other somewhat promising righties take bullpen spots and run with them. That team is no threat to win 100 games, but it could win 87 games.
“But wait”, you say. “Your rosy scenario said nothing about the starting pitching.” Oops. And that is the problem with the 2011 Indians in a nutshell; they might not wind up having a good offense or a good bullpen, but it's not that hard to envision how it might happen. On the other hand, it's very hard to see the starting pitching performing in a manner that could make Cleveland a legitimate threat in the AL Central.
Carlos Santana will be the catcher, and is said to be back to normal after recovering from the knee injury that prematurely ended his rookie season. Even assuming that to be reality, it's easy to see Santana being something of a disappointment to fans in 2011 after hitting .261/.408/.469 in 187 major league PA. Whether those expectations are met or not, Santana should be on of the Tribe's best offensive players and have the bat to carry first base, where he is slated to see some time.
The competition to be his backup will continue throughout the spring. Lou Marson started last season as the Indians catcher, but my money would be on him beginning the season in Columbus. He's still young enough that the organization would presumably like to get him some regular playing time rather than the rare chances that figure to come behind Santana. That leaves an assortment of uninspiring choices (veterans Paul Phillips and Luke Carlin and youngster Juan Apodaca). I would be very surprised if it was Apodaca, and the fact that Marson is already on the 40-man roster may wind up winning him the position after all.
At first base, the hope that Matt LaPorta could be the impact hitter that would really make the CC Sabathia trade payoff have essentially been extinguished--it's now just a question of whether he can be an adequate major league hitter at the position. The jury is still out on that, but his former prospect status still has at least another year or so with which to tease observers. LaPorta has seen time in left field in the past, but appears anchored to first base now.
Second base figures to be a black hole for the team as it has been since Asdrubal Cabrera moved across the bag for good. Luis Valbuena played himself out of Cleveland's plans last season, and while he is in camp and ostensibly competing for the job, the odds against him are staggering. If the team wouldn't promote Santana at the outset of the 2010 season, you can forget about seeing prospect Jason Kipnis (or Lonnie Chisenhall at the hot corner, although he is considered to be farther away anyway). That leaves Jason Donald, who'll likely instead be slotted at third base, and non-roster invitee Adam Everett as options (and let's be honest--if you're not going to use Everett's shortstop glove, you're not going to put him in your everyday lineup).
Thanks to this grim outlook, the Indians decided to bring in a veteran to at least give a steady presence at the position--Orlando Cabrera. Cabrera really doesn't offer much more than his name at this point, but given the choice of watching Luis Valbuena or just about any human being other than Muammar Gaddafi, things could be worse.
Third base is the other position where the team has thrown their hands up. Jayson Nix was thought to be the favorite, and will almost certainly be on the team one way or the other, but his late season trial at the hot corner was scary from a fielding standpoint, and the man has a career 77 OPS+ in over 700 PA, so it's not as if there's a tradeoff being made between fielding and offense. Donald now appears to have the inside track on the position, but he too leaves much to be desired offensively. Jared Goedert, who hit 27 homers between Akron and Columbus last year, will also get a look, but at 26 and without much a previous track record, he's more suspect than prospect. Jack Hannahan is also in camp as an option of last resort.
Cord Phelps, like Chisenhall and Kipnis, could be a midsummer or later addition to the team. Phelps is trained to play some third as Kipnis has surpassed his as the second baseman in waiting, but third base puts him in competition with Chisenhall, so he may really be preparing to take a utility job down the line. The Indians will have decent flexibility, as Cabrera and Donald can both back up short and second; there will be no need for an Everett type to be tacked on the roster solely for Asdrubal Cabrera's days off.
The outfield picture is much clearer than the infield, particularly if Sizemore is able to perform. Right field belongs to Shin-Soo Choo, who is not really appreciated by the home fans anymore than the general consensus. If you tried to tell a typical Tribe fan that Choo was roughly comparable in value to Carl Crawford in 2010, they'd laugh you off as most people from Chicago or San Francisco would as well. Sizemore will likely not be ready for Opening Day; apparently two weeks into the season is the target, but personally I don't expect to see him before May. The brass seems to be committed to playing him in center when he returns, but left is a possibility with Michael Brantley sliding over to center as he will in Sizemore's absence. I remain a Brantley skeptic, as he has yet to display a hint of power; he will need a .350ish OBA to be an asset. His .270 BABIP suggests some bad fortune, but given that his major league OBA is .313 in over 400 PA, it's going to take more than a few hits dropping in to make him a legitimate major league outfielder.
I was not particularly pleased that Austin Kearns was brought back; his hot start obscured the fact that he was a below average hitter in a corner position. He does provide a right-handed bat to platoon with Brantley, and he's borderline playable in center (as is Choo), so he brings a bit more to the table than Shelley Duncan. Duncan is left-handed, more of a liability in the field, and has performed no better at the plate. He'll probably make the team anyway and also see time at DH, particularly on days Hafner is scheduled to sit and a righty is pitching. Travis Buck, signed to a minor league deal, is a more intriguing option, similar but more useful in the field than Jordan Brown (a LF/1B). Both are left-handed, which is problematic given Brantley's presence.
Trevor Crowe would be another candidate for the bench, but he is held back by an injury and may not be ready for some time. Ezequiel Carrera, Chad Huffman, and Nick Weglarz are also in camp but don't figure to win a spot on the team.
Travis Hafner is the DH, and it's worth noting that he still contributes to the offense. 5.7 RG is still an acceptable and even slightly above-average performance from a DH. His production was similar in 2009--5.9 RG, which was actually a little worse relative to the league. The problem with Hafner as a player is that he requires frequent days off due to nagging elbow problems (only 118 games in 2010 and 94 in 2009). The real problem with Hafner as an asset is his massive contract, which still has two years and $26 million to go (one could argue that his platoon splits make him a liability against southpaws, too, I suppose). However, the typical Tribe fan carries on about Hafner the hitter as if he's completely useless, unable to separate judgment of the contract and the memory of the hitter he was from a fair assessment of the hitter he is.
Nick Johnson was a late addition to the first base/DH mix on a minor league deal. Usually I'd be very excited about Johnson joining my team, but with the Tribe not going anywhere there's really no reason to throw obstacles in the path of letting LaPorta take 600 PA. Johnson can also fill in at DH, and gives the Indians the potential to have a number of Secondary Average beasts (joining Sizemore, Choo, and Santana); of course, concerns about his effect on opportunities for other players have a high probability of being rendered moot given Nick the Stick's injury history.
The Indians' starting rotation is pretty well set, barring injury; this can be a good thing if you're the Phillies or the Giants, and a bad thing if you have Mitch Talbot and Carlos Carrasco locked into starting jobs. Fausto Carmona will take the ball on opening day; there has been so much written about him that I feel completely incapable of providing any analytical insight. Carmona stands out from other pitchers that have had flukish Cy Young type seasons (Joe Mays and Esteban Loaiza come to mind in the last decade) because he does have the stuff that makes you want to overlook the low strikeout rates and believe that he could be an exceptional pitcher with an unusual profile. But believing doesn't make it so, and Carmona is a groundball pitcher typecast as an ace by a team desperate for an ace that doesn't have very good infield defense.
Left-handed batters have remained the scourge of Justin Masterson, and as such he might well be best suited for a relief role. However, of all teams Cleveland needs to try to resist that temptation as long as possible. While as a rule I'd usually support not giving up on a pitcher's rotation potential until absolutely necessary, it is even more imperative in the case of Masterson and the Indians. The team is thin in starters, and three of the team's top pitching prospects (Alex White, Jason Knapp, and Nick Hagadone) are considered to be potential future relievers.
Mitch Talbot is a perfectly acceptable back of the rotation filler type who'll likely slot third in this club's rotation because his 159 innings and 14 RAR in 2010 constitute a track record relatively. Carlos Carrasco's stuff has always outranked his results, but he pitched well in September and will thus be given a rotation spot with little resistance this spring.
Fifth starter options are numerous, but most lack potential to ever be much more. Josh Tomlin and Jeanmar Gomez are forgettable righties, David Huff is the last man standing of the Indians once prodigious collection of left-handed soft-tossers with the rade of Aaron Laffey.. Hector Rondon was one of the team's top prospects but has been stopped cold by injuries, while deadline trade swag Zach McAllister and Corey Kluber figure to be in the mix to fill in during the summer, but not break camp with the club. The aforementioned White has more upside and will also be a possibility late in the campaign.
The wildcard in the fifth starter competition is Anthony Reyes, who will be given every opportunity to win the job. Reyes was acquired in 2008 and started 2009 in the rotation before requiring Tommy John. The Indians would love to see him physically able to pitch thanks to his experience if nothing else, but it is pretty clear that he will not be ready for Opening Day. The team was linked to veteran free agents Kevin Millwood, Jeremy Bonderman, and Bartolo Colon; the latter signed with the Yankees, sadly denying Tribe fans of seeing a former contributor return home and perhaps fall flat on his face with a flair not seen in these parts since Juan Gonzalez' one-game appearance in 2005. The former two remain on the market but each day diminishes the likelihood that they will turn up in Cleveland.
Recent developments have made the bullpen picture clearer than one might expect under the circumstances. Chris Perez will be the closer, and is a good bet to disappoint; his .234 BABIP should go up, and even without using DIPS-inspired metrics, his peripheral RA was a good three-quarters of a run above the actual figure. None of this is to say that Perez will be bad, only that one should not expect a repeat performance of the second half, in which he was a lockdown closer.
Rafael Perez and Tony Sipp give the bullpen a pair of left-handers who tease their potential to be more than specialists while often proving combustible on the mound. Chad Durbin, briefly an Indian in 2003-04 (~ 60 innings), has been signed to a major league contract and thus will be the team's primary middle innings reliever. The Durbin signing may seem unnecessary given the club's outlook, but it appears as if giving Manny Acta some hope of a stable, experienced right-handed middle reliever outweighed other considerations.
Those four are locks; two others, Jensen Lewis and Joe Smith, are strong possibilities. Both were signed to contracts rather than being non-tendered, which leads one to believe that they will be pitching in Jacobs Field this year. Smith is a pleasure to watch if nothing else, and Lewis' childhood Indians fandom allows the fans a degree of warm fuzzies.
The final spot appears to be reserved for a long reliever; if that is the case, two pitchers jump to the front: Justin Germano and Joe Martinez. Both have started and relieved in their careers; Germano has the upper hand as his BABIP gave the impression of effectiveness in 35 innings with the Tribe in 2010.
If the team is willing to eschew a long man requirement, Frank Herrmann and Vinnie Pestano will be very much in the mix. Herrmann is a 27 year old righty without good stuff; his performance over 45 big league innings was average but he lived on a low walk rate (1.8 per nine) and a low strikeout rate (4.8). The fall off the high wire could be ugly. Pestano throws hard and impressed in his September callup; his AA and AAA performance was impressive as well (1.81 ERA, 77/16 K/W in 59 innings). It would be nice to see the team go with Pestano, as there really aren't any right-handers with power arms to set up Perez (I don't want to give the impression I have a fetish for power arms--”effective” arms would have worked too).
Other relievers who could appear as the season progresses include Josh Judy, Jess Todd (definitely stuck in neutral at this point), Nick Hagadone (if moved to the pen), Bryce Stowell and Zach Putnam (who'll immediately become my least favorite Indian of all-time). There are a number of arms with potential in that group, and so it's possible that the Tribe bullpen could be a lot more exciting by September and perhaps positioned to be more effective in 2012.
Last year I predicted a 74-88 season, which was five wins too many; I've learned my lesson and will go with 72-90 this year. It's dangerous to attempt to project further down the road, but I could see 2012 as a similar season to 2011 with some more young players breaking in, with 2013 as a season in which to make a push. But by that point the contract status of players like Sizemore and Choo will be an issue, and so it's best to keep an even keel and say that it's not at all clear when Cleveland's next contender will take the field.
A best case scenario (really more like 90th percentile)? Let's say 81-81. The starting pitching makes it tough to go much higher even if one assumes favorable health for the offense and an effective bullpen.
Worst case? Any time you expect to win 72 games, you could lose 100 if things don't go your way. I don't think this is the worst team in the majors, but only a homer could argue that median expectation wouldn't put them in the bottom third or quarter of the league.
Sadly, this is an organization that needs a break from a public relations standpoint. The performance of the organization has been a disappointment over the last three seasons without question, but the fickleness of the fanbase has been a disproportionate response. Indians fans have every right to demand better from the organization, but when just four years ago your team was one win away from the World Series, it is unbecoming to act as fans of a club that has spent a generation planted in the second division. Perhaps more than any other fanbase, Indians rooters have swallowed the vision of baseball as hopeless for small markets hook, line, and sinker--even when they could simply look to the city's other major league franchises and see that salary caps can neither compel “homegrown” players to stay home (even when the team wins a lot of games and the player is actually from the area) or ensure that an organization will put a decent team on the field even once in a decade. The pathetic, feeble whining about a ballclub by denizens of a declining, corrupt city is a bit much for me.
C: Carlos Santana, Luke Carlin
IF: Matt LaPorta, Orlando Cabrera, Jayson Nix, Adsrubal Cabrera, Jason Donald (Shelley Duncan in lieu of Johnson)
OF: Michael Brantley, Grady Sizemore, Shin Soo-Choo, Austin Kearns, Shelley Duncan (Travis Buck in lieu of Sizemore)
DH: Travis Hafner
SP: Fausto Carmona, Justin Masterson, Mitch Talbot, Carlos Carrasco, Josh Tomlin
RP: Chris Perez, Rafael Perez, Tony Sipp, Chad Durbin, Jensen Lewis, Joe Smith, Vinnie Pestano
Tuesday, March 01, 2011
Comments on Bill James Gold Mine 2010, pt. 2
2. Defensive Win Shares and Loss Shares
James has revamped Win Shares over the last couple of years to include Loss Shares. I think this is a very good thing, although I look forward to when (if?) the entire methodology is published. Without the full explanation, it's dangerous to comment about isolated details, but James' essay on "Explaining Defensive Win Shares to a Dead Sportswriter" is tough to ignore. My Twitter-friendly take on it: He's going to have trouble explaining it to a lot of people, not just dead sportswriters.
Again, it's impossible to evaluate the method while knowing so little about it, but James makes this extraordinary statement:
Making outs increases the team's responsibility to play defense. When you make more outs, that increases the team's responsibility to play defense. Therefore, if two players are the same in the field but of them makes more outs, the one who makes fewer outs has to come out ahead when you compare the player's defense contribution to his defensive responsibility.
Lest you think that was just a slip, he doubles down:
While we are in the habit of thinking of offense and defense in baseball as un-connected, they are in fact not un-connected. There is a very important connect between them, which is the rule that for every out you make on offense, you must record an out on defense.
Bill James is obviously a very intelligent man, and you a very intelligent reader, so I am hesitant to respond to this--the response should write itself. Limiting myself to a paragraph or less, I suppose it is technically true that each out on offense is matched by a defensive out, barring walkoffs and rainouts and the like. But there is no causation between the two. The rules of the game require three outs per inning and nine innings per team. Each team makes 27 outs regardless of the rate at which they use them (think OBA) or any other factor.
An individual who makes outs at a higher rate than some comparison player does not increase the number of outs that his defense must record. The defense must record 27 outs regardless of what an individual does at bat. What does happen is that by consuming excess outs, the individual batter leaves less outs to be consumed by the other eight members of his lineup, and fails to generate additional plate appearances for them.
James later seems to suggest that the revamped DWS-LS system assigns the same responsibility to field to each position, regardless of where it stands on the defensive spectrum. He then states his objection to offensive-based positional adjustments, and so it seems as if the stuff about making outs might be a backdoor way of applying positional adjustments. It's unclear, though, and still doesn't follow logically.
James’ discussion of positional adjustments also seems to gloss over the use of defense-based positional adjustments or the fact that most of us who still use offensive positional adjustments do so because we believe they provide a ballpark estimate of the defensive differences between the positions. When I use an offensive positional adjustment, I'm not saying that I think a shortstop with a 5 RG is a better hitter than a first baseman with a 5 RG. What I am saying is that the difference between aggregate offensive performance between shortstops and first baseman (when considered carefully and over a long period of time) approximates the inherent difference in defensive value.
You are certainly free to reject that argument (and many sabermetricians that I respect very much do just that), but please recognize that the sabermetrician using an OPADJ is likely not making the claim that a player's offensive contribution is altered by his fielding position.
More important than my own positional adjustment folly is an apparent failure by James to recognize that the positional adjustments that are now used most prominently in the community (generally Tango's, which have made their mark on the PADJs used in WAR figures from both Chone and Fangraphs) are based on estimates of the defensive difference between positions, sometimes informed by offensive averages. Furthermore, the sources do not lump the positional adjustment into the offensive ranking--they break everything (offense, fielding, baserunning, position, etc.) into smaller components, which are then summed to produce RAA, WAR, or some other total value metric.
Again, it is possible that I have misunderstood James' point, or that he has done a poor job of expressing himself, and that DWS is completely logical. However, I think it is going to take a much more thorough explanation of the system to give people that read the Gold Mine piece a lot of confidence in his methodology.
3. Strikeout rate
One of the most thought-provoking essays is "Whiff 7", which discusses the phenomenon of strikeout rates continuing to reach all-time highs. James argues that there is no end to this in sight under current conditions, as teams have an incentive to find power pitchers but no disincentive to find batters that avoid striking out. James argues that the standard deviation of power (he doesn't use that terminology) has decreased over time, and so league homer rates have gone up while the top individual performers hit about as many homers as they did in previous eras.
James then offers some suggestions of rule changes that would slow or reverse the trend. It's an interesting piece, and it didn't prod me to respond to it directly, but rather to make a tangential and mostly unrelated point about how we measure strikeout rates--a wholly unoriginal and stale one at that.
I have for a long time advocated using K/PA rather than K/IP as the measure of pitcher strikeout proficiency (I’m not claiming this is unique, as others have carried that banner with much more vigor and coherent arguments than I have offered). Through no effort of mine, the use of K/PA has increased in the sabermetric community, with sites like The Hardball Times and Fangraphs prominently utilizing K/PA.
As an example of how the different denominators can change perception, consider the point that most long-term successful pitchers have at least average strikeout rates. This is a point that the average fan still mystifyingly misses a great deal of the time. Take Greg Maddux for example. Maddux is apparently seen by some as a non-strikeout pitcher. Here is a table with his K/9 versus the league average, with KAA being strikeouts above average per inning:

For his career, Maddux struck out 6.1 per nine, while the league average was 6.4. He struck out 206 less batters than an average pitcher would have in the same number of innings. Without seeing the same figure for a lot of pitchers, it's hard to contextualize that, admittedly.
Suppose that instead you look at Maddux through K/PA:

Now Maddux' strikeout rate is essentially average--he struck out 17% of opposing batters, the same as the league average. Maddux' career rate is lower (it's actually 16.5% to 16.6%), but just barely so, and by this metric he only recorded 22 less strikeouts than average.
In Maddux' peak years (I think 1992-98 stand out), he was above-average even by K/9--+90 KAA, while he was an even more robust +196 when K/PA is the standard.
This is not intended to recast Maddux as a strikeout fiend--certainly he was not, even at his best. Still, Maddux' strikeout rate is more impressive when viewed in light of the number of opposing batters he actually faced rather than in terms of innings pitched, which really is just a measure of the percentage of outs a pitcher gets via the K rather (this is obscured by displaying strikeouts per 9 innings rather than strikeouts per 27 outs).
In addition to K/9, there are several other per-inning pitching ratios in common usage--H/9, W/9, HR/9, WHIP. What all of those have in common is that they are ratios of bad things (offensive successes) to good things (outs recorded). K/9 is a ratio of really good things (outs recorded by strikeout) to another set of plain old good things that includes the really good things (total outs recorded). As such, it's best viewed as a measure of a pitcher's reliance on strikeouts.