Accountability
How good is this model, actually?
Every rating, projection and probability on this site comes out of one model. This page marks it against games it was never shown, and prints what it got wrong.
12.6 pts
Average miss
on the final margin
69.9%
Winner called
of held-out games
15.7 pts
Typical error
root-mean-squared
See every prediction, game by game →What it says about this week →
Measured on 2025 games the model never saw: it is trained through an early week, then scored on everything after. Missing the final margin by 12.6 points on average is not precision — it is roughly two scores, and it is why nothing here is presented as a prediction.
Is it better than doing nothing?
A model is only worth having if it beats the obvious alternatives. Both are scored on exactly the same held-out games.
This model
12.6 pts · 69.9% winners
Opponent-adjusted ratings
Unadjusted margins
13.0 pts · 67.9% winners
Each team rated by its own average margin, ignoring who it played
Always the home team
17.2 pts · 55.2% winners
Home side by the home-field constant, every game
Adjusting for opponents is worth about 0.35 points of accuracy over not bothering — real, but smaller than you might expect. Both comfortably beat picking the home team every week.
When it says 70%, does 70% happen?
Accuracy only asks whether the pick was right. Calibration asks the harder question: whether the model's confidence is honest. Every held-out game is sorted by how sure the model was, and compared with what actually happened.
| Model said | Games | Predicted | Actually won | Miss |
|---|---|---|---|---|
| 50–60% | 232 | 55.1% | 55.2% | +0.1 |
| 60–70% | 213 | 64.7% | 57.3% | -7.4 |
| 70–80% | 163 | 74.9% | 73.0% | -1.8 |
| 80–90% | 128 | 84.3% | 83.6% | -0.7 |
| 90–100% | 132 | 94.8% | 98.5% | +3.7 |
This table used to look much worse. The model originally converted a margin into a probability using its average margin error (15.7 points). That answers a different question from “how often does the favorite win”, and it made the model systematically too timid: it said 84.7% for games favorites actually won 98.0% of the time.
Two things were wrong at once. The rating engine deliberately compresses extremes — a margin cap and shrinkage toward zero — so a real mismatch arrives with its gap already squashed. And margin error is inflated by blowouts, games whose winner was never in doubt. Both push the same way.
The fix was to stop assuming the relationship and measure it: a one-parameter logistic regression of “did the favorite win” on predicted margin, fitted on 868 held-out games. It puts the effective spread at 11.0 points instead of 15.7. Average calibration miss across the bands fell from 4.5 points to 2.8 points.
Checked out of sample, not just in. Fitting on half the held-out games and scoring the other half gives the same answer, and the spread lands within about half a point whichever half is used — so this is a property of the data, not a curve bent to fit it.
What is still off: the 60–70% band now misses by -7.4 points in the other direction, on 213 games. One slope cannot bend the curve band by band, and at that sample size the gap is roughly two standard errors — so it may be noise rather than bias. It is left visible rather than smoothed away.
Show the old, uncalibrated numbers
| Band | Games | Said | Won | Miss |
|---|---|---|---|---|
| 50–60% | 327 | 55.1% | 53.5% | -1.5 |
| 60–70% | 243 | 64.6% | 67.5% | +2.9 |
| 70–80% | 154 | 74.5% | 81.2% | +6.6 |
| 80–90% | 99 | 84.7% | 98.0% | +13.3 |
| 90–100% | 45 | 93.1% | 100.0% | +6.9 |
Every test, separately
Three train/test splits, so a single lucky cut cannot flatter the result.
Trained through week 8
396 games · 12.5 pts · 68.2%
Trained through week 10
291 games · 12.5 pts · 71.8%
Trained through week 12
182 games · 12.8 pts · 69.8%
How it is tested
The model is fitted on games up to a cut-off week and then scored on every game after it, which it has never seen. That is repeated at three different cut-offs and the results averaged, so no single split can flatter the outcome.
Probabilities come from a spread fitted on held-out games — a one-parameter logistic regression of whether the favorite won on the predicted margin. Bands with fewer than fifteen games are dropped rather than shown, because a win rate from a handful of games says nothing.
Games with no favorite, and games involving a team outside the ratings pool, are excluded — there is nothing to be right or wrong about in the first case and no rating to test in the second.