How the method would have done, season by season, before it went live. Walk-forward: every game predicted using only games played before it. · predictions are published to show the model works. nothing here is betting advice.
65.0%
straight-up, 4,146 regular-season games 2010–2025
66.5%
Vegas closing favorite, same games
±1.5
95% margin on that accuracy (points)
0.219
Brier score · Vegas 0.211 (lower is better)
87%
of picks agree with Vegas
10.4
avg. margin error (pts)
What this says, plainly: over sixteen seasons the model picks the winner about 65 times in 100 without ever seeing a betting line; the market's favorite wins about 67 in 100. That gap is real but small, and no public model closes it entirely — the NFL is a coin with a heavy side, not a solved game. The useful claims are the calibration below (a 70% pick really does win about 70% of the time) and that the live record is locked before kickoff.
Season by season
Regular season only. One season is ~270 games, so ±6 points is normal noise around the long-run average.
Bars: our walk-forward accuracy. Dashed marks: the Vegas closing favorite the same season.
Season
Games
Our accuracy
Vegas favorite
Brier
2010
240
61.7%
66.7%
0.234
2011
256
68.4%
66.8%
0.212
2012
255
63.1%
64.3%
0.221
2013
255
63.5%
71.4%
0.216
2014
255
69.4%
66.7%
0.214
2015
256
64.1%
62.5%
0.226
2016
254
65.0%
64.4%
0.219
2017
256
65.6%
70.8%
0.208
2018
254
68.9%
66.1%
0.212
2019
255
64.3%
64.3%
0.224
2020
255
67.1%
67.5%
0.213
2021
271
61.3%
62.4%
0.227
2022
269
62.5%
66.2%
0.222
2023
272
63.6%
68.0%
0.227
2024
272
70.2%
71.3%
0.205
2025
271
60.9%
65.3%
0.223
When we differ from Vegas
The honest cut. On the 3,613 games where our pick and the closing favorite agree, both are right 68.1% of the time. The 529 games where we differ (13% of games) are where the whole gap lives.
43.9%
our pick wins, when we differ (529 games)
56.1%
Vegas favorite wins, same games
41.5%
when we back the home team against Vegas (301)
46.9%
when we back the away team against Vegas (228)
Read this the right way round: a game where we differ from the market is not a hidden edge — it is usually a game where the market knows something the public feeds don't. The live version of this record is on the Record page.
Calibration over the backtest
Stated confidence bucket vs. actual hit rate. Bars should climb in step with the labels.
90–100% (42)
95%
80–90% (394)
82%
75–80% (376)
76%
70–75% (560)
72%
65–70% (578)
67%
60–65% (693)
60%
55–60% (752)
57%
50–55% (751)
55%
How it was run
The rule book lives in docs/BACKTEST_METHOD.md.
For every week from 2010 to 2025, the model was trained on every completed game before that week (history starts in 2009) and asked to pick that week's games. The same feature builder and model code make the live picks, so the backtest cannot be better than live by construction. A leak audit rebuilds features for sampled games from strictly-prior data and asserts they match. Ties are pushed. The Vegas column uses the closing spread from the nflverse schedule feed; the model never sees it.
Known limit: the historical starting quarterback is the one who actually started; live picks use the named starter from the feed at run time, so a late scratch can cost a live pick that the backtest would have got.