Showing posts with label Picking a Starter. Show all posts
Showing posts with label Picking a Starter. Show all posts

Tuesday, April 5, 2011

FIP: A thing that is better than ERA at predicting future performance



This is a continuation of previous posts looking at how predictive different pitching metrics are. To estimate this, I plotted the relationship between a stat in year one versus ERA in year two, and calculated the R^2 value for the line of best fit - for a detailed description of the method, check here. The R^2 value for ERA in year 1 vs ERA in year 2 is 1.3 - therefore for a stat to be better at predicting future ERA than ERA itself, the R^2 value must be > 1.3. The stats analysed are those most commonly referred to by fantasy baseball tipsters to suggest which pitchers you should pick up because they are good, and who you should avoid because they are just lucky.

If you are interested in K:BB ratio, check here

EDIT and here is the first of these articles from 2011. They use xFIP to assess pitcher performance over an incredibly small sample size. Read on to find out if that is a good idea.

If you want to know what is better than ERA at predicting future ERA (e.g. an R^2 > 1.3) read on:

Whilst FIP may be not very fair to the 'lucky' pitcher when calculating WAR, it is the best metric we have to estimate future pitching performance. Check the graph.

Whereas the other metric beloved of fangraphs loving fantasy baseball experts, xFIP, is barely better than ERA at estimating future performance:

So, in summary: If you want to use stats to guess a pitchers future performance, use FIP, not ERA. But it isn't perfect - an R^2 of 0.2 isn't that high. Furthermore, it's actually quite easy to out-perform FIP - in the next graph for each qualified starter I've calculated ERA-FIP, and then plotted the average for all qualified starters by year. In this case, a negative number = outperforming FIP - as you can see, on average a starter will outperform (have a lower ERA) than FIP.





And finally here is a table of pitchers who've been qualified starters each of the last 3 years, showing how often they've had an ERA outperforming FIP.




What does this mean? Well, those pitchers who've had a better ERA than FIP for the last 3 years will probably carry on outperforming their peripherals and those who have a better FIP than ERA will probably continue to under-perform their peripherals - but this doesn't matter too much as Tim Lincecum and Justin Verlander are still pretty good, even if they aren't as excellent as they 'should' be. The bottom line is that FIP is the best stat for predicting future performance, so you should be interested in pitchers who's FIP is better than their ERA. But it isn't perfect, and there are many pitchers who will have better or worse results than their FIP suggests.

Wednesday, March 9, 2011

How to Pick a Starter Part I

There will, in the upcoming baseball season, be many, many blogs trying to identify those pitchers who've been unlucky (e.g. James Shields), and should be acquired at all costs, and those pitchers who've got great numbers which are little more than luck (e.g. Trevor Cahill), and therefore must be traded ASAP to some poor sap who doesn't follow advanced statistics.

So which stats should we use to judge future pitcher performance? Can stats divide the good from the lucky and the unlucky from the rubbish? To keep things simple, we'll begin with looking at what past ERA tell us about future ERA.

The Method

From Fangraphs, I downloaded the pitching stats of all qualified starters going back to 2005. In this initial analysis, I asked how consistent a pitcher's ERA is from year to year. To do this, I plotted for each pitcher his ERA  in year 1 vs his ERA in year 2 - and to increase sample size (n=267), I repeated this for each pair of years - for instance, in today's graph, the data is the correlation between 2005 ERA and 2006 ERA, 2006 ERA and 2007 ERA etc, all the way through to 2010. I've treated each yearly pair as an individual data point so that many pitchers stats appear more than once. By looking at the correlation between all these individual pitchers' ERAs from year to year, we can judge how good ERA is at predicting future success (at least as measured by ERA).

How good is this year's ERA at predicting next year's ERA?


We can measure the correlation between data points by calculating the R^2 for a line of best fit through the data points. If every pitcher's ERA was identical from year to year, R^2 would equal 1. Obviously this is impossible. Instead, the (low) R^2 of 0.13  for the correlation between this* year's ERA and next* year's ERA suggests that an individual pitcher's ERA is very variable from year to year.

Helpfully, this data also gives a baseline, so we can ask whether a given pitching stat (e.g. WHIP, K:BB ratio, FIP, xFIP) is better (R^2 >0.13) or worse (R^2<0.13) than ERA at predicting future ERA and in this way identify the stat that is the best predictor of future pitching success). Those exciting analyses are coming very soon.

*where this and next can be any pair of consecutive years.