Showing posts with label Draft Strategy. Show all posts
Showing posts with label Draft Strategy. Show all posts

xWHIP 2.0: The Next Generation

The following article is from my most recent article for The Hardball Times.

A few months ago, I debuted the first version of the expected WHIP (xWHIP) calculator, which took a pitcher's batted ball distribution and, in determining an expected number of hits, calculated that pitcher's expected WHIP. The tool was tinkered with and refined until version 1.4.3 was released and that, until now, has been the primary xWHIP tool available. xWHIP 1.4.3 overexpected WHIP a bit, but was otherwise pretty solid. Especially for relative comparison purposes, xWHIP 1.4.3 was a useful fantasy tool.

Not long ago, I was introduced to a fellow stathead by the name of Martin Alex Hambrick. He had done some number tinkering similar to what I had done independently with the xWHIP calculator, and he had an idea. He brought that idea to my attention, and from it a new formula for expected xWHIP was born.

Alex's idea was that a pitcher's actual innings pitched (aIP) are as much the by product of luck as expected hits (xHits). The theory is that a medley of defense, umpires, errors, random luck and the like skew the length of innings. The pitcher, for example, does not particularly control dropped third strikes by his catcher.

This idea is somewhat captured in the K% (K/TBF) and BB% (BB/TBF) movement of sabermetrics that rejects K/9 and BB/9 because the length of innings is largely out of the control of the pitcher, thereby skewing both K/9 and BB/9. Accordingly, we began work on a new denominator for xWHIP that incorporated an expected innings (xIP) total based on a pitcher's outs-creating events.

With this idea in mind, we began work on a new xWHIP calculation. Law school delayed my work on a final formula until this week, but with "way too much time on my hands" (i.e., any lawyers out there need a law clerk for the summer?), I finally got around to hammering out a reliable formula and user-friendly interface, calibrated to Baseball Info Solutions (BIS) data.

The current formulation for expected innings is as follows:

xIP = ((K*1.000075)+((BB-IBB+HBP)*0.00016)+((0.808)*GB)+((0.278)*LD)+((0.992)*IFFB)+((0.745)*OFFB)+(0.020099*(BB+HBP+xH)))/3

The coefficients in the above formula represent the expected outs by event rate. You might notice the two percent adjustment applied to both modified walks (BB-IBB+HBP) and expected hits (xH). That figure represents a ten-year average outs-per-runners-put-on-base rate (ORB). ORB encapsulates the ten-year league average pickoff and caught stealing rates.

Because catcher defense and a pitcher's pickoff talents are difficult to measure, and also not widely available, using a league average rate helps make the calculator more accessible. The final xWHIP figure should be mentally modified based on one's own perception of a catcher's pickoff ability or a pitcher's pickoff ability. If Jason Varitek is the catcher, you might want to raise the pitcher's calculated xWHIP, while the opposite would be true for those pitchers handled by Yadier Molina.

Alex is working on a simplified "Quick xWHIP" formula that simplifies the xWHIP calculation even further, to the point that you could do it on a calculator. He'll tell you more about that (and the accuracies of both xWHIP 2.0 and Quick xWHIP) in a (near-) future post. All I can say for now regarding the calculator's accuracy, at least to some degree of certainty, is two things. First, xWHIP works best—that is to say, it is most predictive—when you use multi-year data rather than year N-1 data. Second, the R^2 of the data seems to be solid for a predictive state.

Someone once told me (or maybe I just read it somewhere) that an R^2 of .30-.35 is strong for a predictive stat, while a .60 or greater R^2 is what is required of an evaluative stat. Using 2007 xWHIP 2.0 to predict 2008 actual WHIP resulted in an R^2 of .34 amongst the 78 pitchers who faced a minimum of 500 batters, compared to an R^2 of .26 for 2007 actual WHIP. Likewise, using 2008 xWHIP to predict 2009 actual WHIP resulted in an R^2 of .36 amongst the 80 pitchers who faced a minimum of 500 batters, compared to an R^2 of .30 for 2008 actual WHIP.

Strangely, however, using 2009 xWHIP to predict 2010 xWHIP amongst the 82 pitchers who accrued 500+ total batters faced merely resulted in an R^2 of .15 (compared to a .14 R^2 for 2009 actual WHIP). Maybe I crunched the 2009-2010 data incorrectly. Maybe this is a sample size issue. Maybe not. As I mentioned above, Alex will supply more details on the accuracy of xWHIP 2.0 shortly.

I also tinkered some with the expected hits formula, but the changes are relatively minor and hardly warrant discussion. The important thing to note about the new xWHIP tool is that it is now calibrated per the past five years of BIS data rather than Game Day. I have done this because I believe that Fangraphs utilizes BIS, not Game Day, as their source for ball in play (BIP) data. Accordingly, this should make the tool more accurate for the average user. Most of the data stood relatively stable, but here are the new expected hits by batted ball types:
  • Popups: .004
  • Groundballs: 0.236
  • Outfield Flyballs: 0.250
  • Line Drives: 0.716
These data points include home runs, which is why the Outfield Flyball expected hits rate is so high. If you take home runs out of the equation and account for them separately (as the xWHIP calculator does), the expected hits rate, per BIS, for Outfield Flyballs and Line Drives falls to .158 and .714, respectively.

You can download the new xWHIP tool, version 2.0, by clicking here. The password to utilize the xWHIP tool is still "soto 18" and the batted ball data you will need to plug in can be found at Fangraphs.com.

Picture below is a screenshot of the xWHIP 2.0 tool, which was used in my Zack Greinke forecast article. For explanatory purposes, this screenshot has the 2010 numbers of Roy Halladay plugged in.


As the instructions on the tool indicate, the gray cells are for data you should manually input. The magenta park factor cell is also a manual data cell, though the number should be left at "1.00000" unless you have the relevant park factor HR/FB index figure. You should not enter any data into any of the blue, green or yellow-orange cells.

The green cells feature the line drive-regressed expected-ball-in-play data. The yellow-orange cells display the expected innings, expected hits and expected WHIP for the pitcher, irrespective of defense. If you enter data into the Team Innings Pitched and Team UZR gray cells, then the blue cells will display a crude defensive adjustment to the expected hits total, assuming uniform defense and that all saved hits would be of the singles variety. All of the data cells are pre-formatted to visually round all numbers to keep the sheet clean, though cells will retain the full value of any number entered.

I also included a cell for xWHIP 1.4.3, calibrated from Game Day to BIS, in case people wanted to know a player's expected WHIP using expected hits and actual innings, rather than expected innings.

I hope everyone enjoys this. If you have any questions/concerns/comments/criticisms, please post them in the comments below or email them to gameofinchesblog@gmail.com, with the subject line "xWHIP 2.0 Calculator."

On a final note, I would like to give a special thank you to several of my THT colleagues who have been invaluable in the creation of the xWHIP 2.0 tool. Without the assistance of Derek Carty, Dave Studemund, and Harry Pavlidis, none of this would have been possible. I apologize to each of you for my incessant e-mailing in attempt to work out the mathematical kinks in the formula.

Fantasy Outlook: DME's Ultra Sexy Fantasy Baseball Resources Revealed

As promised, my two money leagues have drafted and I have a few interesting files to share with you, the public. You can navigate the files below by the labeled tabs at the bottom of the file. Some of the files are .xlsx files and will require Office 2007 or higher (though they can be converted to .xls files for free).
  • Draft Comparison Rankings - this first file compares Mock Draft Central (MDC) rankings to those of Yahoo fantasy sports (Y!) and ESPN.com (ESPN). The theory behind this file is that MDC rankings represent free market player valuations (courtesy of mid-February data). By comparing MDC rankings to those of Y! and ESPN, one finds who is under/overvalued in each service. Nelson Cruz seems to be perfectly ranked in both Y! and ESPN!, while Y! seems to underrate Max Scherzer and ESPN seems to undervalue Dan Uggla. EDITOR'S NOTE: Yahoo Sports updated its player rankings on March 11, 2010. I have updated the data to reflect these changes.
  • Starting Pitchers Rankings - this second file has my Top 50 Value SP rankings (players are ranked based on draft position value; thus, no Tim Lincecum or Roy Halladay) and a fully sortable spreadsheet listing all the important peripheral and 2009 statistic data for all 130 starting pitchers who logged 100+ innings last season (data courtesy of Fangraphs).
  • xBABIP-based 2010 hitter projections - Using 2009 data for all MLB hitters who amassed 300+ PAs last season, I projected 2010 triple slash lines for 284 hitters (assuming that all hits added/subtracted due to "luck" were singles). For more information on how the data was calculated, click here.
A few other valuable resources I utilized for my draft and auction leagues:
I wish everyone the best of luck in fantasy this year. I will post some more analysis in the coming days. Long live King Felix.

_______________________________________________

Want more Game Of Inches? Click here to follow Game Of Inches on Twitter.

Sexy Rexy's Drafting Strategy

So I have helped you (or like two people who may have actually read and liked my drafting tips) on some basic tips on how how to draft and maintain a good team. If you talked to me in past week or so, you know I have told you some ideas on my thoughts on players but I did not want to divulge too much information because there are some keys guys I wanted and did not want others to really know about them. So here we go,

1) I literally did research on every single projected starter. This allowed me to see which positions had the most depth. To me, 3B, 1B, SP, and OF had the most depth. I thought 2B and especially SS were awesome at the top and while I thought they had little depth but still some quality guys like Kelly Johnson, Rickie Weeks, Christain Guzman, and Ryan Theriot that were available in much later rounds. For me, if I didn't get J-Roll, HanRam, or Jose Reyes in the first round, there was no reason to draft a SS until like the 18th round, which allowed for great flexibility. Despite some quality 2B I could get around the 15th round or so, I still wanted a top tier 2B which I had to get in the top 3 rounds.

2) I took on the Sam Walker/Fantasyland approach of the trifecta pitching. This really isn't hard to do in non-auction leagues but I really wanted some stud pitching at the top of my "rotation". Yes, there is plenty of SP depth, but if I stocked my team with late round SP, my "rotation" would be nothing more than average. Because I was able to get a quality SS and quality OF later in the draft, I was able to get those positions I wanted later while getting above average to great SP in earlier rounds. A big problem I found is that a lot of top tier pitchers were young and abused/ in the WBC so seemingly more prone to injury, guys like: Tim Lincecum, Cole Hamels, Jake Peavy, and Roy Oswalt immediately came off my list. This left the top tier pool mainly consisting of Arizona pitchers. I ended up taking CC Sabathia as my first pick and although he's abused and I didn't like it a whole lot, I feel his past and his age make him a better candidate than most other pitchers.

3) My ideal players are guys with high BA (above .280), can go at least 15/15, and a high OPS. (To me, because R and RBI are not stats based on individual performance- they're based on the team and position in lineup- if a guy has power and can get on base a lot, they're more inclined to get R and RBI). Surprisingly, there were plenty of these guys which were: Ian Kinsler, Brandon Phillips, David Wright, Hanley Rameriz, Jose Reyes, Ryan Braun, Alfonso Soriano, Grady Sizemore, Carlos Beltran, Hunter Pence, Matt Holliday, Jayson Werth, Elijah Dukes, Nate McLouth, and Matt Kemp. Pretty elite group, no? The problem is that once the 3rd round starts, you'd be hard pressed to find any of these players. But two guys on this list stuck out to me: Jayson Werth and Elijah Dukes. These are two quality guys that go under the radar and thus I can get at lower rounds. So then, all I did was take a closer look between these players to determine which ones I wanted.

4) There still are certain positions that needed to be filled so I filled those positions with specialized players- mainly players with a lot of HR or SB, preferably with a BA and OPS if possible. I used these players to fill 1B, OF, 3B, and Util (mainly depending on my position in the draft, i.e. if I got Pujols or Wright I wouldn't need to fill 1B/3B respectively) The big name that jumped out to me was Jacoby Ellsbury- a guy who will probably by one the of the best OF in terms of SB with a high BA and a decently high OPS. For 1B, I needed pure power foremost then comes BA. Because I was just decent in HR everywhere else, I just targeted Carlos Pena. Although his BA pisses me off a bit, I think it's OK considering my high BA elsewhere and like I said, HR comes first. The 3B I targeted was Garrett Atkins. In my research, he came up as an Aramis Rameriz but with a .300+ B.A., which comes extremely useful to me. And because he was less values in the draft than other 3B like ARam, I could get a quality 3B a bit later in the draft.

5) So here was a list of specific players I WANTED: Kinsler, Ellsbury, Dukes, and Werth. I love Kinsler and I knew that I probably couldn't get him unless I had a late first round pick, and just my luck, I had the 9th overall pick. In the first three rounds I wanted a combination of 2B, SP, SS, and OF. (Now I clearly would have to change my strategy depending on on my draft position but...) I was able to get Kinsler in the first. Now I wanted Johan Santana in the 2nd and thus get the best available OF in the third, but Since Johan was taken I went with one of the best OF on the board that fit #2 (which was Soriano- admittedly though I took this pick in haste) and luckily CC came to me in the 3rd. Honestly, I thought CC would get taken so the next top tier SP would have been Brandon Webb (and maybe in some ways Webb would have been better) but I didn't want to reach and Sabathia was still available. Now base on previous drafting I knew Ellsbury generally went in the 4th and Atkins generally in the 5th so while they may have been slight stretches, I wanted those players. So in the first five rounds I got a 2B, Of, SP, OF, and 3B. I knew Werth generally went in the 12th so that's where I picked him up and Dukes and Guzman/Theriot (really the only SS I liked outside the first round) go really late so I was able to get some quality SP and RP, I think, between rounds 8-15.

7) I had back up plans just in case things went wrong. Joey Votto was a great 1B to pick up in case the 1B went higher than positions I wanted to take a 1B; Corey Hart was my substitute to Ellsbury; I could get a quality 2B like Brian Roberts, Dustin Pedroia, or Brandon Phillips in the 3rd round or Rickie Weeks and Kelly Johnson in much later rounds; I could have (and did) gotten a SS really late in the draft; 3B was deep so if someone overvalued Atkins I was willing to go after I guy like Edwin Encarnacion in much later rounds; Chris Young would have been an acceptable substitute for Werth; Jack Cust and Jim Thome are fine power/Util guys in later rounds... Anyway, despite my structured attempts in round 1-6, I had a back up just in case.

So now I have given up my secrets for this year's draft. The only draft that mattered was David "MVP" Eckstein's pay league and now that that's done, feel free to take my secrets. Plus, my strategy will change for the next drafts, I could care less really if I lose others leagues, and I still have some players up my sleeve. Enjoy picking at my brain!

Saves

In fantasy baseball, saves are overrated. With picks like Papelbon early, players pay high for 10% (in a 5x5 standard) of their impact points. When four category players like Curtis Granderson are still on the board, it is silly to waste a pick on a guy who will surely rack up saves and slightly help your ratios when you know that late in the draft there will be no 3/4 category players left, buts still plenty of guys who will get you saves and not hurt your ratios.

In the large scheme of things, the average RP tosses what, 70 innings? Most players employ 2-3 closers, depending on both the quality and availability of closers in the draft/free agency pool. Even with 3 guys, the average total impact of those relievers would be 210 IP, the innings equivalent of what a top 30 starter would toss. Assuming an 1800 IP limit that come standard in a Yahoo roto league, that cumulative 210 IP comprises less than 12% of your team's total innings allocation. A combined RP core with a 3.00 ERA, 180 K, 210 IP, 1.10 WHIP line (with some scattered Ws here and there) would slightly help a team whose starting pitching core were to average a 3.75 ERA, 165 K, 1.3 WHIP and 190 IP per SP (based on what it takes to win according to Roto Authority) by boosting some average ratios, but how much would adding an inferior RP core that racked up a comparable amount of saves across the same IP hurt a team?

Let's say instead of drafting Papelbon, Mariano Rivera and Bobby Jenks (a combined 120 SVs) last year, you drafted Brian Wilson, Kevin Gregg and George Sherrill (a combined 102 SVs)? You would not necessarily clear the saves category for 10 points, but you would still be in a prime position to rack up respectable saves numbers with late round draft picks. While last year Papelbon (49 ADP), Rivera (72 ADP) and Jenks (107 ADP) combined for a valuable 2.11 ERA (47 ER), 14 W, 120 SV, 192 K, .90 WHIP line across roughly 202 IP, Gregg (292 ADP), Sherrill (237 ADP) and Wilson (155 ADP) combined for a 4.41 (90 ER), 13 W, 102 SV, 183 K, 1.40 WHIP in about 184 IP. While the latter groups line isn't nearly as pretty as the elite RP's line, the variance between the two groups is much more overstated that you'd expect, based on 2009 ADP.

quick sidenote: obviously some changes to this grouping would need to be made on the basis of facts like Sherrill being replaced by Chris Ray midseason or Kevin Gregg not being the closer (even though he should, allowing the Cubs to continue to maximize Marmol's use in high leverage situations), but these guys could easily be substituted for late round 2009 closers like Chad Qualls (204 ADP) or Joel Hanrahan (198 ADP); the general point of this argument remains the same.

As you might observe, the difference in counting stats productions between these two groups of RPs in 2008 was somewhat negigible. 9 Ks and 1 W is more the byproduct of circumstance than opportunity. As noted earlier, being 18 SVs short of 3 of the top 5 closers going into 2008 with late round picks would put you in prime position to stay competitive in the SVs category.

The combined line of the elite RP core would surely improve your teams ratios to better approach "what it takes to win", but at the cost of 3/4 category offensive guys. On the other hand, the 4.40 ERA over 184 IP accumulated by the inferior RP core would constitute a 10.2% impact on all 1800 IP, causing a 3.75 ERA line to rise to 3.81. By contrast, the elite RP core's impact would constitute 11.2% of all IP and lower ERA from 3.75 to a 3.56 line. A similar impact is observable on WHIP (the elite RP core lowers whip to 1.25, the inferior RP core increases WHIP to 1.31).

What we can observe here is that while the elite RP core does positively impact pitching statistics, the inferior RP core simultaneously NEGLIGIBLY impacts the pitching line -- while the superior RP core lowers ERA well below the projected "what it takes to win" threshold (while also lowering WHIP a sizeable chunk), the inferior RP core only increases ERA by .06 and whip by .01. With smart drafting of SPs (who have much more impact on cumulative pitching statistics than RPs, mind you), this tiny impact of the inferior RP core could easily be offset, while gaining the benefit of a 3/4 category hitter or two early in the draft (by forgoing the elite RP in favor of an "inferior one").

What is very roughly observable here is the overrated impact of RPs based on draft position. Inferior RPs can easily put up comparable counting stats without hurting your ratios if you draft smart. Comparatively, late game hitters (minus sleepers, which can be harder to effectively forecast) generally cannot produce comparable offensive numbers to early round hitters. Come picks 180+, the hitters left in the pool are generally one category guys and offensive gambles. If you go with an early RP, late hitter strategy, you are maximizing risk (by forgoing more reliable/valuable hitters) and minimizing overall return, which is somewhat irrational in the investment game known as drafting. Smarter investment would call for the strategy of simply ignoring RPs until the later rounds (or just ignoring saves all together and focusing on the other 9 categories; why not try maximizing 90% of your potential against balancing 100%? It is equally as viable if done properly).

A lot of this knowledge is "conventional wisdom" for baseball drafters; I'm just putting some rough numbers behind the assertion. Now that that is established, let's move on to the point of my post: rules of thumb for drafting Closers (btw, I should go back in my post to modify all references to RPs as CLs, but you can infer what I mean and as I mentioned earlier, I'm lazy).

If it isn't obvious by now, I think closers are overrated. I rarely draft them (because so many guys lose their jobs midseason), but when I do, I like to look for bargains. I look for guys who meet the following criteria (in this order):
1) Job security
2) Save opportunity potential

First and foremost, if you are going to draft a closer, you want to draft a guy who is going to close. Why waste a pick on a guess (ie, who is closing in Seattle for 2009) when you can probably pick up the person who the closer out of spring training loses the job to off of waivers within a week or two? Why also waste a pick on a closer who has a better RP who may unseat him from the closing role behind him? To me, it seems like a waste of a pick that you could spend gambling on a guy like Denard Span or Shin-Soo Choo.

Secondly, I like a guy who is going to get the chance to save games. This doesn't mean a closer for an offensive friendly team like the Red Sox or Yankees, but closers for teams like The Pirates and Royals, where every one of their 70 wins per season is a save opportunity. Last year, I got Matt Capps, Joakim Soria and Brian Wilson at incredibly discounted prices and they paid off big time.

If you want my recommendation of who to draft, I'd recommend Brandon Lyon (because no one in Detroit is really healthier than him who can throw strikes), Brian Wilson (who has great stuff and no one to unseat him), Heath Bell (the Padres are notoriously committed to keeping their closer the closer during the season, no matter how hard he struggles) and Grant Balfour (the healthier, younger, better option for Tampa; Balfour for Closer is as inevitable as was Obama for President).

That's all for now; it's late and I'm tired. Apologies for spelling/grammar errors; I just wrote this from beginning to end in one sitting and am (yes, you guessed it) too lazy to re-read/edit tonight.