How it works
Sources: ProFootballTalk, CBS and ESPN feeds; the National Weather Service forecast at kickoff for outdoor stadiums; the nflverse injury report and named starters are what the model already counts. The reader (one language-model call per game, only when there is new news) returns typed claims about things not already counted. A fixed table turns each claim into points, scaled by the reader's confidence, capped per type and ±10 in total. Weather is logged at zero weight until a season of grading says otherwise.
| Claim type | Max points |
| starting QB change not yet in the feed | 6.0 |
| key starter out, not on the injury report | 2.0 |
| key starter returning | 1.5 |
| resting starters / playing backups | 8.0 |
| suspension | 2.0 |
| coaching change | 2.0 |
| motivation / eliminated / nothing to play for | 1.5 |
| locker-room or organizational turmoil | 1.0 |
| unusual travel or schedule disruption | 1.0 |
| weather at kickoff (logged, zero weight) | 0.0 |
| other (in-week, concrete, high confidence only) | 0.5 |
Scorecard by claim type
After each final: did the move go the right way, and did it lower the squared error (Brier)? Negative Brier change is good. Types that don't earn their weight get zeroed at season's end.
| Type | Claims | Graded | Avg pts | Helped | Brier Δ |
| key starter out, not on the injury report | 7 | 0 | 1.9 |
– | – |
| starting QB change not yet in the feed | 1 | 0 | 5.7 |
– | – |
| other (in-week, concrete, high confidence only) | 1 | 0 | 0.4 |
– | – |
Shadow scorecard — what we declined
Rejected claims are kept at zero weight and graded as if applied. If a rejection reason keeps "helping", the rule is too strict; if "other" keeps hurting, it stays out. This is how the layer learns.
| Rejected because | Type | Claims | Graded | Would have helped | Brier Δ if applied |
| confidence 0.65 below floor for other | other | 4 | 0 |
– | – |
| retired: now on the injury report | key_player_out | 4 | 0 |
– | – |
| confidence 0.75 below floor for other | other | 3 | 0 |
– | – |
| confidence 0.70 below floor for other | other | 2 | 0 |
– | – |
| no ruled-out language (practice status is not 'out') | key_player_out | 2 | 0 |
– | – |
| not an in-week coaching change | coaching_change | 2 | 0 |
– | – |
| already on the injury report | key_player_out | 1 | 0 |
– | – |
| confidence 0.00 below floor for coaching_change | coaching_change | 1 | 0 |
– | – |
| confidence 0.30 below floor for key_player_out | key_player_out | 1 | 0 |
– | – |
| confidence 0.35 below floor for key_player_out | key_player_out | 1 | 0 |
– | – |
| confidence 0.60 below floor for coaching_change | coaching_change | 1 | 0 |
– | – |
| no ruled-out language (practice status is not 'out') | qb_change | 1 | 0 |
– | – |
| return of a player the model was not counting out | key_player_return | 1 | 0 |
– | – |