Showing posts with label DVOA. Show all posts
Showing posts with label DVOA. Show all posts

Tuesday, August 7, 2007

Do Stats Need to be Adjusted for Conference?

In this post and this post, I've found that the proportion of games won by the home team and those won by the "better" team (according to DVOA) is smaller for interconference games with a larger variance in the former proportion than for intraconference games. Games between divisional opponents have the highest of both proportions and the smallest variance in the former proportion. This led me to hypothesize that coaches are able to better adapt to teams they see more frequently, which makes inuitive sense, but I was also curious whether or not stats such as DVOA (or my VOLA stats) needed to be adjusted based on conference averages rather than league averages. Are NFC teams rewarded unfairly by DVOA for playing three quarters of its season in a conference full of weak teams? Ideally, opponent adjustments will filter out those effects, even if it's by league average. But perhaps 4 games against a stronger conference is not enough to filter out the effect. NFC Team X has played Y% better than league average when adjusted for opponent, but they were more likely to because of an easy schedule. Against AFC opponents, they'd be less likely to play at that level, even when adjusted for opponent strength. This effect might pop up more in games between average to below average opponents rather than when the NFC team is legitimately good. If the effect is real, then we'd see a greater proportion of interconference games being lost by the team with the higher DVOA when that team is in the NFC rather than the AFC.


I went back and looked at games from the entire history of DVOA (1996-2006). The DVOA totals used reflect the entire season's performance, including postseason. We should see the balance of power shift in favor of the AFC in 2000-6, given that only 1 NFC team ('02 Bucs) won the Super Bowl in that time span, and that team had a large and rare tactical advantage.

















Year% Was Better Team% Was Better Team
and At Home
Better Team Win %
Overall
AFCNFCAFCNFCAFCNFC
19960.566670.433330.283330.216670.676470.65385
19970.566670.433330.250.183330.558820.53846
19980.550.450.283330.233330.696970.7037
19990.650.350.316670.166670.743590.57143
20000.650.350.316670.166670.641030.7619
20010.50.50.250.250.80.8
20020.593750.406250.296880.203130.631580.57692
20030.609380.3750.281250.156250.666670.66667
20040.750.250.406250.156250.770830.5625
20050.656250.343750.328130.171880.619050.63636
20060.593750.406250.296880.203130.710530.5



The first table shows that the AFC might have been the better conference overall since at least 1996. For both conferences, the number of games in which their team was the better team is split evenly between home and away in almost every year, so neither conference's numbers should be unduly affected by home field advantage. For 6 out of 11 years, the AFC had a better win percentage when they were the better team than the NFC did when they were. In 2 of those seasons, the difference amounted to less than one game ('96 and '97). Looking at 2002-6, after the realignment, the AFC has had the better win percentage in 3 out of the 5 seasons. In 2003 and 2005, they were essentially even. In 2004, the NFC would have had to have won 3.33 more games in which they had the better team to match the AFC's win percentage. In 2006, they would have needed 5.4737 games. The predictive power of DVOA (where prediction = team with better total DVOA wins) was only 55.56% in 2006, compared to 66.67% in 2005 and 65% in 2004. Based on the table above, I'm guessing a large amount of the dropoff was because of interconference games, valuing NFC teams too highly given the gap between the conferences. That the year-to-year differences in better team winning percentage favor the AFC more strongly on average also indicates NFC teams being rewarded by a weak conference.

















YearBetter Team Win %
Home Better
Better Team Win %
Away Better
AFCNFCAFCNFC
19960.764710.769230.588240.53846
19970.60.818180.526320.33333
19980.823530.785710.56250.61538
19990.789470.70.70.45455
20000.842110.80.450.72727
20010.80.866670.80.73333
20020.789470.769230.473680.38462
20030.833330.90.523810.5
20040.807690.60.727270.5
20050.76190.818180.476190.45455
20060.736840.461540.684210.53846



The NFC teams see a bigger gap in better team win % for home and away games. The average difference is 18.517% for the AFC and 22.807% for the NFC. Small sample size has some effect for the games in which the NFC team is better, however. On average, the difference seems to be one or two games that could go either way.

So DVOA does seem to unfairly reward some NFC teams for playing in a weaker conference. Which teams are throwing off the system, though: the good teams being bumped up to very good or the mediocre teams being bumped up to good? In most years, the average gap in DVOA between teams in games in which the AFC was the better team and lost was larger than that for the NFC. And in most years, the average DVOA of the better teams that lost was higher for the AFC than the NFC. I would say, then, that the NFC teams that are mediocre to slightly above average (Rams, Falcons, Panthers) are unfairly rewarded more than good NFC teams (Eagles, Cowboys) for playing in a weaker conference.

Though DVOA is far more sophisticated than my methods, I believe I can apply the lessons learned here to my prediction model, specifically my VOLA stats. For intraconference games, I might start using Value Over Conference Average rather than VOLA. For interconference games, I could use Value Over Other Conference Average. So if Team A average 6.5 yards per play, it could be average in the AFC, but it might above average in the weaker NFC. Perhaps you could do a similar thing for DVOA, but the math of it is rather hazy to me right now.

Read More......

Monday, August 6, 2007

How Often Does the Better Team Win?

In the National Football Leagues, no win can be guaranteed. One play can have a large impact on any game. It's part luck, but it's part happenstance too. It takes skill to get an interception, but it's the happenstance of what play is being run, and thus where the players are, that largely dictate if the interception is returned for a touchdown or not. If the interception occurred in the red zone and was returned for a touchdown, that's at least a 10 point swing. So a play with a 2.96% probability of occurring (league average interception rate) has an inordinately large impact on the game. Brian Burke's blog had a very good piece on how much luck is involved with winning and concluded that half of winning games is luck (52.5% to be exact) so the better team is going to win around 74% of the time. I was curious to see how it worked out in reality and further validate my assertion that interconference games have more inherent variance and less predictability than intraconference games. To decide the better team, I simply used total DVOA from Football Outsiders (1996-2006). Please note that the DVOA stats are over the entire regular season and postseason, so they are retrodictive, not predictive. The predictive ability of total DVOA is not as good. If DVOA had a predictive accuracy of 70%, I wouldn't be working as hard on a prediction system.



Average result means how many more points the better team scores on average. Average margin of victory means how many more points the winning team scores on average. The averages are by year, so the total proportion of games won by the better team, etc. will vary slightly from the numbers listed here. Part of the original study on interconference study was to see how much year-to-year variance there was in the outcomes of those games.









InterconferenceInterdivision
Better Team Win %Avg. ResultAvg. Margin of VictoryBetter Team Win %Avg. ResultAvg. Margin of Victory
Mean, 1996-20010.680566.922212.2280.664465.69411.042
Std. Dev, 1996-20010.0798732.50310.997760.0477331.5790.75167
Mean, 2002-60.656.931312.2130.689586.539611.169
Std. Dev, 2002-60.0450151.15370.899530.0485241.49970.92161












IntradivisionAll Games
Better Team Win %Avg. ResultAvg. Margin of VictoryBetter Team Win %Avg. ResultAvg. Margin of Victory
Mean, 1996-20010.69066.444911.2320.680966.356911.42
Std. Dev, 1996-20010.0475671.11090.705740.026391.03430.24626
Mean, 2002-60.733337.072911.3560.696096.837511.5
Std. Dev, 2002-60.0502811.0890.530040.0275380.648520.31041



So with the 20/20 hindsight of each entire season, the better team has wins about 69% of games, close to the 74% reported in Brian's blog, which was based on 2002-6. When looking at 2002-6 intradivision games, he was pretty much dead on. 73.333% vs. 74%. Most of the discrepancy can be traced back to interconference games. The divisional realignment in 2002 reduced year-to-year variance in the percentage of games won by the better team, but it's also reduced the average percentage from 68% to 65%. It's interesting that the average margin of victory is larger in interconference games than in the other types, but I'm not sure what that means.

I have two ideas on possible reasons why fewer interconference games are won by the better team. First, maybe coaches have more problems adapting strategy to opponents they don't see as often. An interconference matchup occurs only once every four years now (before, some matchups were much more common than others). Coaches have only a week to prepare for games, so they can only learn so much about a team's strengths and weaknesses. Obviously, the more time they have to study opponents, the more they will learn about them. So every time the interconference matchup comes up, the coach probably has to throw out a good deal of what he learned the last time. With intradivision matchups, you see the opponent twice a year and can re-use knowledge gained from previous matchups. Second, maybe stats should be adjusted for conference quality in addition to specific opponent quality like in baseball. I'm not sure this would work, given that the rules in both conferences are the same, unlike in baseball. Given that 75% of the season is intraconference, though, perhaps it's slightly inaccurate to judge a team based on the whole league, rather than their specific conference, when trying to predict an intraconference game. I've toyed with implementing this idea and might pursue it sometime in the near future.

Read More......

Sunday, August 5, 2007

Another Look at the Importance of Offense and Defense in the Playoffs

In this post, I examined why defensive performance has a higher correlation with playoff success and determined that the root was the greater abundance of very good offenses among playoff teams. Thus, teams needed better defenses to make it through the playoffs. It's not a matter of one unit being more important than the other. It's a matter of balance. In this article, I redid the previous experiment using Football Outsiders' DVOA stats instead of my own VOLA stats. DVOA covers the same time period as the original experiment: 1996-2006. It's very important to point out that my VOLA statistics covered only the regular season, but the DVOA stats used here are based on regular season and postseason (where applicable) performance. This probably biases certain stats used here, such as the average DVOAs of Super Bowl winners, simply because the DVOAs used take into account that the teams performed well against good teams in the postseason, which isn't necessarily indicative of their regular season performance (see the 2006 Colts). The results, however, support what I said the last time.


For the following tables, O=Offense, D=Defense, R=Run, P=Pass, ST=Special Teams.

The correlation coefficients with the seeds are actually the correlation coefficients with 7-(Seed #), where the seed # is 7 for non-playoff teams. So the second column includes all teams (where the seed input is between 0 and 7), but the third column includes only playoff teams (where the seed input is between 1 and 7).












Corr To WinsCorr to Seed (all teams)Corr to Seed (playoff teams only)
O0.64780.495780.19407
RO0.46970.357970.23226
PO0.62410.479160.11531
D-0.5252-0.3899-0.27018
RD-0.3517-0.25361-0.202
PD-0.5066-0.37634-0.25634
ST0.26030.237870.18571



For this table, defensive DVOAs less than or equal to -X% are tallied in the ≥X% columns.











ALL TEAMS≥10 (Teams)≥10 (%)% Made Playoffs≥20 (Teams)≥20 (%)% Made Playoffs
O6819.82579.412216.122490.476
RO4513.1266.667113.20763.636
PO10129.44673.2674914.28685.714
D7722.44968.831174.956382.353
RD9728.2853.608216.122447.619
PD9527.69761.053349.912570.588
ST10.291550000














PLAYOFF TEAMSMedianMean
O6.657.475
RO1.551.1735
PO11.8513.123
D-7.15-6.903
RD-7.9-7.0205
PD-8.55-6.8765
ST0.950.97121














SB WINNERSMedianMean
O11.512.582
RO63.9273
PO17.820.936
D-15.1-14.155
RD-12-10.955
PD-14.9-17.373
ST2.22.3455



Teams with very good offenses are more common than teams with very good defenses, and teams with very good offenses are more likely to reach the postseason than teams with very good defenses. Having either still gives you a pretty good chance of making the playoffs. Having both makes it nearly certain. 28 out of 30 teams (93.33%) that had at least a 6% Off. DVOA and at least a -6% Def. DVOA made the playoffs. The only 2 teams not to were the 2002 Miami Dolphins, who blew a big lead late in the final game of the season against the Pats, and the 1999 Oakland Raiders, who were #3 in total DVOA but posted an 8-8 record.

Defense, however, has the higher correlation with playoff seeding. When nearly 70% of playoff games are won by the home team, seeding becomes very important. Then again, a higher seeding usually indicates a better team. Looking at the DVOAs, the home team is the better team in about 67% of postseason games, and the home team wins 69% of postseason games. The home teams win 70% of games in which they are better and 60.6% of games in which they are worse (though this number has dropped to 40% since the realignment in 2002).

On average, the Super Bowl winner is a balanced team. The offense DVOA is above 10%, and the defense DVOA is below -10%. And actually, the average defense DVOA comes out slightly higher. Special teams DVOA is positive as well. The worst offensive DVOA of a Super Bowl winner belongs to the 2000 Ravens (-7%), and the worse defensive DVOA belongs to the 2006 Colts (11%). The Ravens, however, at least had a good running game (8.3%). In 11 Super Bowls, the team with the better total DVOA won 9 times. The two exceptions were the 2001 Patriots and the 2006 Colts. The 2006 Colts, however, had the better weighted DVOA, in which the latest games are weighted more strongly. Seven champions had better offensive DVOA, and seven champions had better defensive DVOA. Six champions had better special teams DVOA. Four had better offensive and defensive DVOA. Only 2 champions were better in all three categories. All of the champions were better in at least one category (2001 Patriots had better ST DVOA than the Rams).

People often point out the Ravens and the Bucs to say that defense wins championships. While Baltimore won in 2000 with a -7% offensive DVOA, they had a better total DVOA than the Giants, who had a good defense (-8.2%) but a weak offense (4.4%) compared to the Ravens' monstrous -30% defensive DVOA. The Ravens also had better special teams (7.2% vs -4.4%). Similarly, the Bucs had better special teams and defense, which was historically good. The Raiders' offense, which was damned good (24.4%) couldn't hold up, but it's hard to quantify the impact of Barrett Robbins and Jon Gruden's vast insider knowledge. The first paragraph of the game summary section of Super Bowl XXXVII's Wikipedia entry indicates that the Raiders' offense was at a huge disadvantage because of Gruden's knowledge. So you can win with a mediocre offense (4 out of 11 had <1% DVOA), but it's reliant on the matchup (the Raiders, the 2003 Patriots' 0.4% Off. DVOA and the Panthers' -7.2% Off. DVOA) and you probably need a ridiculously good defense. Only two teams, however, have won with a mediocre defense (>0% DVOA). The 1998 Broncos had a really good offense led by John Elway and Terrell Davis (28% DVOA), and the 2006 Colts had a really good offense. The Broncos played a well-balanced Falcons unit (10.1% Off, -13.4% Def, 2.8% ST), too.

Teams like the '00 Ravens and '02 Bucs have struggled since their Super Bowl wins because defensive performance is less consistent on a yearly basis, and they did not have the capability on offense to make up for the performance loss. The Colts will struggle to duplicate their success in '06 as well without at least an average defense.

Read More......

Tuesday, July 10, 2007

Creating Power Rankings

Just as a lark, I figured I'd try devising some power rankings using my model. The basic idea behind them is this: If every team played each other once at home, once at away, who would win the most?

I tested this idea on the 2006 teams after the end of the regular season. Using my adj. rush, adj. pass, sack, adj. 3rd down conv., and turnover data, I trained a linear regression predictor on all games before 2006. Then I used the model to predict the outcomes of all 992 possible games. I created 2 separate rankings: one based on the sum of margins of victory/defeat, one based on the expected winning percentage.






































RankTeamSum of Exp. Margins
1PHI4.2027
2SD4.1368
3BAL3.8437
4JAX3.2917
5DAL2.1643
6CHI2.1514
7PIT1.9942
8NO1.8954
9IND1.7482
10NE1.5237
11MIN1.1935
12CAR1.1847
13ATL0.94039
14NYG0.84285
15MIA0.82191
16GB0.71736
17CIN0.52797
18DEN0.33802
19KC0.29031
20SF-0.685
21WAS-1.0359
22STL-1.4769
23NYJ-1.5889
24SEA-2.0378
25BUF-2.3402
26TEN-2.4492
27TB-2.8237
28ARI-3.2092
29DET-3.6094
30OAK-3.7756
31HOU-3.8764
32CLE-4.9007







































RankExp. Win %Team
1tBAL0.83871
1tPHI0.83871
3SD0.80645
4JAX0.72581
5tCHI0.66129
5tNO0.66129
7NE0.64516
8tDAL0.62903
8tCAR0.62903
10tPIT0.6129
10tIND0.6129
10tMIN0.6129
13tMIA0.59677
13tATL0.59677
15NYG0.58065
16GB0.56452
17tCIN0.54839
17tDEN0.54839
17tKC0.54839
20tSTL0.43548
20tSF0.43548
22tNYJ0.41935
22tWAS0.41935
24SEA0.33871
25TEN0.32258
26BUF0.27419
27tOAK0.19355
27tARI0.19355
29tCLE0.17742
29tHOU0.17742
29tDET0.17742
29tTB0.17742


Looking at the bottom, it's about what you'd expect. Seattle, a playoff team, shows up at #24 in both rankings, but they were 8-8 in a weak division and suffered some losses due to injury and free agency. The Jets, another playoff team, show up at #23 and #22 respectively due to weak rushing offense and defense and average passing offense and defense. Seattle and the Jets were #25 and #19 respectively in Football Outsider's 2006 DVOA rankings. Philadelphia, despite a 10-6 record, shows up at #1 with a very strong offense (14.721% VOLA rushing, 20.58% passing unadj.). Jacksonville also obtains a very high ranking (#4) despite its record, a mediocre 8-8, thanks to well above average rush offense and defense. New England is a little low compared to DVOA (#10 and #7, #5) with average and below average YPA numbers but high sack and interception rates. The Super Bowl champion Colts are at #9 and #10 (DVOA #7) with 20.142% VOLA in pass offense and 26.09% VOLA in pass defense(!too high!) but -28.29% VOLA rush defense.

So overall, the rankings turn out very similar to the DVOA rankings with individual teams gaining or losing a couple spots. There are no major disparities, however, that I can see, except maybe Dallas showing up too high in the ranking by points. It would be interesting to calculate these rankings over the season and compare these with other rankings.

Figures and article corrected on 7/12/2007 after finding errors in some box scores.

Read More......

Saturday, June 16, 2007

The Value of Home Field Advantage Part I

So let's go back to basics. Really basic stuff.

How much is home field advantage actually worth? Pretty straightforward question, but there are several ways to tackle the question. From 1994-2006, home field advantage was worth about 2.6362 points on average, but the standard deviation of the results was about 14.0780 points. 69.91% of the games fell within one standard deviation of the mean, or in other words, those games were within 2 touchdowns either way of the average result. As a side note, the outcomes fall in line with a normal distribution, which means that linear regression is a good way to try predicting future outcomes. So home field advantage on average matters to a small extent. Home teams win 58.81% of games. Between two teams very close in talent, go for the home team, but if one team is clearly better than the other, you're better off picking the better team. That's pretty much the conventional wisdom, isn't it?


So how does a predictor like the spread do in terms of valuing home field advantage? The average spread from 1998-2006 was -2.5346, very close to the actual average. The standard deviation, however, is only 5.7783. The extreme outcomes (games decided by 20+ points) have a 18.515% chance of occurring. There's very little incentive statistically to predict such large wins, though the outcome is more frequent than is perhaps expected. I ran a few experiments to classify games as big wins or close wins and for which team in the original research, and the more classes I introduced, the worse classification accuracy became. For the 4-class problem, 28-32% accuracy was the best I could do. The prediction systems play the odds and thus have a tighter range of margins than the actual outcomes. Tightening the bounds of the actual range within the training data does not help accuracy of the prediction systems I've implemented.

What's interesting to note is that the value of home field advantage fluctuates a fair deal from year to year, but reached a peak in 2005 and a deep, deep valley in 2006, which caused the accuracy of the spread and my prediction systems to similarly fluctuate, particularly on 2006. The chart below lays out all the specific numbers.














Average Actual ResultAverage Spread*Proportion of games won by home teamProportion of games home team was favorite
OVERALL** 2.6362 -2.534658.51%66.859%
19983.5042-2.447962.917%67.083%
19993.0645-2.391159.677%64.516%
2000***2.8226-2.614957.447%68.511%
20012.0444-2.264155.645%65.323%
20022.2461-2.263757.813%64.844%
20033.5313-2.558661.328%70.313%
20042.5078-2.537156.641%66.406%
20053.6484-2.630958.984%67.969%
20060.84766-2.802753.125%66.797%















Average Actual Margin of VictoryAverage Margin of Victory Predicted by SpreadProportion of games won by favoriteProportion of games in which favorite beat the spread
OVERALL**11.4715.349865.352%48.263%
199811.5045.731370.00%52.917%
199911.3555.556565.726%50.403%
2000***11.7985.823464.682%45.957%
200111.0775.159365.323%49.194%
200211.1054.935562.109%49.219%
200311.9144.979267.13%51.389%
200411.3675.12762.891%44.922%
200511.6885.41672.656%48.047%
200611.4265.521558.594%46.094%


* Spread is negative when home team is favored.
** Spread covers 1998-2006, but the averages for actual outcomes are from 1994-2006.
*** Spreads for week 4 of 2000 could not be found and are not included in the 2000 spread stats.


Curiously, there's a correlation with the predictive performance of Football Outsider's DVOA stats as well. In 2005, two-thirds of games were won by the team with the higher DVOA. In 2006, that number fell to 55.80%. Without more years of data, it's hard to say if this is just some natural aberration. But there was something unusual about 2006. Was it rule changes, a change in how rules are enforced, a change in stadiums or playing fields? Is there anyway to account for the natural variance from year to year? For prediction systems like linear regression, one could alter the bias coefficient to reduce the bias towards home teams, but there's no guarantee that it'll improve accuracy.

In what ways are the spread and other prediction systems being inefficient in dealing with home-field advantage? One obvious place to start is the weather.

Read More......

Previous Works

This is an updated version of the "Previous Works" page from my research website:

Neural network quarterbacking: Michael Purucker

Michael Purucker began researching the use of neural networks for predicting NFL games in 1996. He used 5 basic statistics based on each team’s performance over the previous 3 weeks: Total yards gained – total yards allowed, Rush yards gained – rush yards allowed, turnover margin (takeaways – giveaways through fumbles and interceptions), time of possession, and victories. Out of all the networks used on the problem, a back propagation network performed the best, achieving a 70.83% accuracy rate over weeks 14 and 15 of the 1994 season. Using the Las Vegas spread, which predicts the winner and the margin of victory, the results improved to 75%. Two weeks, however, is a very small test set. In the 2 weeks, the BP network with the spread was 9 of 14 and 12 of 14 respectively. Three games is a very big variance, and there’s nothing to guarantee it won’t go 6 of 14 in some weeks. The rush yards statistic overlaps with the total yards statistic, and both of the statistics are comparing a team’s offense with its own defense by subtracting yards allowed from yards gained. This does not reflect the matchups that actually take place on the field. The victory input does not take into account the margin of victory, and close games can come down to random events that would have let either team win.


Neural Network Prediction of NFL Games

Joshua Kahn continued Purucker’s study, testing statistics from the entire season in addition to the previous 3 weeks and eliminating the victories input. Kahn cites Purucker’s system as being 60.7% accurate over time. Kahn’s 3 week averages were 37.5 and 62.5% accurate over weeks 14 and 15 of the 2003 season respectively. Using season-long averages, Kahn achieved 75% accuracy in each of those two weeks, while the ESPN experts achieved 57% and 87% accuracy on average (~72% 2-week average). The study has the same problem of an extremely limited test set, but the results demonstrate an interesting point. Using season-long averages rather than 3-week averages yields a better predictor. This makes sense given the larger sample sizes involved. Teams that start off the season 0-3 almost never make the playoffs, so a 3 game winning streak that leads to a 3-10 record does not reflect the team’s quality. That the ESPN experts experienced a 30% variance in the 2 weeks, like the predictor using 3-week averages, could reflect how recent performance can skew human perception.

NFL Point-Spread Ratings

Roger Johnson uses only the Las Vegas spreads to formulate rankings for each team in the league and based on those rankings, predicts the winner of each game for weeks 3-17. For 2003-2006, the system has averaged about 64% accuracy, making it slightly less efficient than the simple “Las Vegas favorite wins” predictor. About 65% of the teams favored by Las Vegas in weeks 3-17 of the 1999-2006 seasons won.

NFL Computer Handicapper Home Page

Daniel Imamura’s “Computer Handicapper” takes more advantage of the data found in the box scores to produce efficiency ratings for various aspects of each team. Rather than just yards per game, many of the ratings are based on yards per play but also factor in turnovers and touchdowns. Using these metrics along with the Las Vegas spread, the handicapper predicts the winner and the margin of victory. Though primarily intended for use on betting with or against the spread, the system has been 55-68% accurate over the 2001-2006 seasons in simply predicting winners.


Football Outsiers

FootballOutsiders.com has gone beyond box scores and into play-by-play data and their own game charting project to devise an array of new statistics, the centerpiece being Defense-adjusted Value over Average (DVOA). The idea behind DVOA is that yards per game statistics are lossy data because the amount of yards gained in an individual play varies in true value based on the context of the down, yardage to go, field position, time left in the game, and the current score margin. In DVOA, each play is categorized as a success or failure and assigned some value based on the context. Given the baseline rates of success and the success value for each play’s context, a team’s overall performance is assigned a value over average, which is then adjusted for the opponent’s average performance. DVOA can be broken down into performance in any situation and by a certain subset of players, allowing for a very fine-grained evaluation of why a team is likely to win. The coarse team total DVOA, however, has shown to be a good predictor as well. In weeks 3-16 of the 2004 season, the team with the higher total DVOA won 67.3% of games. After week 17, the accuracy fell down to 65.625%, having predicted 7 out of 16 correct games. In the last week of the season, many playoff positions have already been decided, so some teams will rest their starters, which could skew the results. In 2005, the team with the greater DVOA won 66.67% of games, but in 2006, accuracy plummeted to 55.80%.

Two Minute Warning: Using NFL injury reports to predict winners

Injury reports categorize players as probable (P(Playing)=75%), questionable (50%), doubtful (25%), or out (0%). Using 1-P(Playing), an injury score is assigned to each player, and the score for each team is the sum of injury scores for all of its players. The team with the lower injury score won 52% of games in 2001-4. Teams with injury scores of 3 fewer points than opponents won 55% of games in that same time frame. When looking at changes in injury score from week to week, the team with the lower injury score "delta" won 55% of games. My original research included injury score inputs. In 2001-6, the team with the lower injury score won 51.947% of games (week 3-17 only). In the same time frame, the team with the lower delta won 52.036% of games. The obvious problem with the injury score is that it's not weighted for player values, but it's not entirely clear how to best do that.

Read More......

Friday, June 8, 2007

The Initial Research

The website built for the original research. Contains graphs, tables, and more complete descriptions.

Rather than predict the exact final score of games, which is dependent on a good deal of random factors, I tried to create a system that could predict the margin of victory/defeat for the home team. The idea being that the better team will win and the better they are, the more points they'll win by. Again, random factors do play a part in the final score margin, but it at least reduces the number of possible outcomes to approximately 93. From 1994-2006, the maximum margin of victory was 49 points, and the maximum margin of defeat was 43 points. Essentially, given what we know about the two teams, I'm trying to see what the expected margin of victory/defeat, the average of the possible outcomes weighted by their probability, is.

Using box scores from this site, I gathered statistics I would use as inputs: rush offense vs. defense (yards per game), pass offense vs. defense (YPG), punt return vs. return coverage (yards per return), sack rate made vs. sack rates allowed (sacks per pass play), time of possession (season average), and turnover ratio (season cumulative). Rather than look at each metric individually, I wanted to compare a team's offense with its opponent's defense in an input to keep in line with the output of "how much better do I expect this team to be." Football Outsiders being a major inspiration for the research, I didn't want to use YPG statistics, but I was forced to. To capture some of what they did, I took the statistics (except TOP and turnovers) and turned them into values over league average.

Value over league average = (Team average - League average) / League average

For defensive measurements, where giving up fewer yards than league average is desirable, I simply multiplied VOLA by -1.

In a similar fashion, I tried adjusting for opponent quality, weighting a team's performance in each game by the opponent's performance over league average. If the opponents gain 150 yards per game rushing and the league average is 100 yards, holding them to 150 yards rushing is an average performance rather than below average. Each yard allowed is worth only 100/150 = 2/3 of a yard. Thus...

Adjusted VOLA = (Team's adjusted average - League average) / League average

If the equations are just stating the obvious, my apologies. To compare home team rush offense with away team rush defense, I merely subtracted the away team's rush defense AVOLA from the home team's rush offense VOLA. Same with the converse and for pass offense, pass protection, and punt returns.

In addition to the statistics, I also used home field climate as an input. Based on Football Outsiders' work, I divided home field climate into 4 types: warm, cold, dome, and Denver (high altitude). For each team, there are 4 binary (0 or 1/true or false) variables representing the 4 possible climate types.

So we have all of these statistics. What do we do with them? Well, as I was studying artificial intelligence when I did this, I went for methods such as artificial neural networks and support vector machines. The methods were tested on the years 2000-2006, one at a time, and training on all years previous to the test set year. I also tried comparing my results against the spread, which similarly represents the expected margin of victory/defeat.

All methods, including the spread, had the following weaknesses:


  1. 10 point barrier: Predictions were conistently off by at least 10 points on average from year to year. Only a couple methods on a couple years could get to a mean absolute error of 9.7-9.8 points. Interestingly, games in 1994-2006 were won by 11.329 points on average.
  2. Too many games classified as wins for home team: 58.51% of games were won by the home team from 1994-2006. Methods would regularly classify 65-80% of games as home team wins.
  3. Small range of predictions: The predictions usually ranged from about -10 to about 15. The actual range of outcomes is -43 to 49. 27.681% of games from 1994 to 2006 were won by more than 15 points. My guess is that this inflated the error, though I haven't checked this out for sure. As soon as I typed this, I put it on the to-do list, however. If you graph the actual outcomes vs. the predicted outcomes, it ends up looking like a parallelogram as seen below.
  4. Sensitive to yearly variations in home-field advantage: The average outcome of an NFL game from 1994 to 2006 was the home team winning by 2.63 points. 58.51% of those games were won by the home team. In 2006, however, only 53.125% of games were won by the home team, and the average result was the home team winning by 0.84 points. As a result, the performance of all methods, including the spread, was aberrantly poor. In 2005, the average result was the home team winning by 3.677 points, and the home team won 59.14% of games. As a result, the performance of all methods was aberrantly good.




The methods I used ran into the following problems:

  1. Some predictions close to zero: About 10-15% of the predictions were that the game would be won by less than one point. Furthermore, about half (5-8%) of those predictions were that the game would be won by less than half a point. As long as they're still not zero, they can still be used for win/loss prediction, but it just doesn't look good to say "Team A is predicted to win by 0.15 points." I'd be curious to see what the win/loss accuracy of these predictions is.
  2. Input stats have very low correlation to margin of victory/defeat: The problem with YPG stats is that they don't necessarily represent quality. Teams with comfortable leads run the ball more to eat up the clock, so the rushing metrics have the highest correlation to the margin. Correlation does not equal causation, however. It's possible the team used the passing game to rack up points quickly, and the defense took care of the rest. Conversely, teams that are behind will abandon the running game in favor of the passing game, which covers more ground in less time. Thus, the passing metrics have a low correlation to the margin.

    Other highly correlated inputs were the turnover ratios and the sack rate metrics.


Specific numbers are given here, but overall, the spread was clearly the best predictor in terms of the following metrics:


  • Win/loss accuracy
  • Average error (how many points off was it?)
  • Correlation of predictions with actual result
  • Proportion of games classified as home team wins.


Using the spread as an extra input, I could get better results in one or two areas in most years, but the spread was clearly carrying the load. Support vector machines with the spread did the best overall, giving stable predictions (unlike neural networks), but they're not easily interpretable models. Linear regression without the spread was 57-63% accurate, which was on the lower end of performance, but it's a model that is easily interpretable.

Where to go from here
For interpretability issues, I'm going to be mainly experimenting with linear regression. The tradeoff in accuracy isn't worth it at this point.

If the spread's the best predictor, then I think trying to model the spread could yield some knowledge that leads toward better predictions. It's a simple matter of replacing the final score margin with the spread as the output I'm trying to predict. I'll be following up on this soon.

The bias towards the home team in all of the methods needs to be taken down a notch. In the case of linear regression, 3.2-3.6 points were being added in favor of the home team automatically (via bias term), which is up to a point more than what should be added on average. Rather than using the bias linear regression comes back with, I could use the actual average result for that year, so in 2006, when the average result was 0.84 points in favor of the home team, an extra 2 points wouldn't be added. The bias could also be adjusted for the home field climate type, eliminating the need for the binary variables.

Most importantly, better inputs are needed. All of the computing power in the world isn't going to help otherwise. From the box scores, I can also include kickoff returns and third down conversion rates. I'm also going to follow up on that soon. Other than that, I think something like Football Outsiders' DVOA statistics are necessary. The DVOA stats take the context of every play into account, filter out random noise, are adjusted for opponent quality, and break down into very specific situations and for specific players and groups of players. DVOA for the pass defense actually measures quality of the pass defense, unlike the YPG stats.

As long as this article is, I glossed over a good deal in terms of specific results. To put it shortly, we could do better. In a way, I spent several months discovering what I pretty much knew already: YPG stats aren't very valuable, warm-weather teams have trouble in cold weather. But it was nevertheless interesting to quantify things like home-field advantage (more on this coming). The research is going to need time to evolve, and my resources are limited. Can't be afraid to fail.

Read More......