Our NFL Model's Real Record: Every Prediction, Graded, Published

Here's a number you'll almost never see a prediction service publish: 52.8%.

That's our machine-learning model's lifetime accuracy against the spread — 821 correct, 733 incorrect, 43 pushes across more than 1,500 graded predictions, every one of them stored in our database and displayed on the site, hits and misses alike. In 2025 it went 148–134–3 (52.5%).

If you're used to services claiming 65% winners, that number probably looks underwhelming. That's exactly why we're leading with it.

What 52.8% actually means

The point spread exists to split every game into a 50/50 proposition. It's a crowd-sourced forecast, sharpened by enormous amounts of information, and decades of research show closing spreads are among the most accurate public predictions of game outcomes in any sport. Against a benchmark engineered to be a coin flip, sustained accuracy above 50% is the entire signal — and models that hold even the mid-50s against closing spreads over large samples are genuinely rare.

So 52.8% over 1,554 graded predictions is the interesting part — but be careful how much weight it carries. That record sits about 2.2 standard errors above 50%: suggestive, not settled. Our own hold-out testing puts the model's honest out-of-sample accuracy at roughly 50% (more on that below), so the fair summary is that we may have a slim edge and cannot yet prove it. Anyone quoting you a sustained 60%+ against the spread is describing either a short hot streak or a fiction.

What we thought at first — and what changed our mind

Our first accuracy figure was 58% against the spread, and we believed it for a while. It was also suspiciously good for a market this well-calibrated, and that is what made us go back and ask what the number was actually measuring.

It turned out to be the selection score — the best accuracy found while searching thousands of feature combinations for a model to promote — not a result from data the search had never seen. A score like that is optimistic by construction: you are reading off the winner of a search, and some of what made it win was luck that will not repeat. Ours had been quietly doing the work of a hold-out number without ever being one.

So we rebuilt the evaluation around a three-way split — train on 2003–2021, select on 2022–2024, and hold 2025–2026 back as a test set the search never touches. Then we ran four independent searches and scored each winner once, on those untouched seasons:

Run Selection score Test score (untouched)
ATS #1 58.7% 50.7%
ATS #2 58.5% 47.9%
ATS #3 57.5% 48.9%
ATS #4 52.8%

Average test accuracy: ~50.1%. The gap between selection and test was 8–9 points on every single run, with different feature sets each time. Which side of 50% a run lands on is search-draw luck, not signal.

Honest out-of-sample accuracy against the spread is therefore about 50% — coin-flip, against a benchmark engineered to be a coin flip. We retrained the final model through 2024 to rule out stale training as the cause; it made no difference, which is what an efficient market looks like.

That leaves the live record as the more interesting number, not the less: 52.8% over 1,554 graded predictions. It is measured on games nobody could tune against, because they had not been played yet.

The week-by-week truth: variance is violent

The 2025 season, week by week, is the best variance education we can offer. The same model, the same process, every week:

  • Week 5: 13–1. The kind of week that gets screenshotted and shared.
  • Weeks 2–3: 21–10 combined. A month in, the model looked unstoppable.
  • Weeks 14–18: 32–45. A five-week 41.6% slump — from the same model that went 13–1 in October.
  • Final: 148–134–3.

Any single week of a forecaster's record is marketing. A five-week window is still mostly noise. The season-long number is where signal begins, and the multi-season number is the only one that matters. This is also why we publish the weekly breakdown on the site — the ugly stretches are the proof the good ones aren't curated.

Picking winners: a higher number against a higher bar

Straight-up — just predicting the winner, ignoring the spread — our model has run at 63% (85–50) since we began grading that objective in Week 11 of 2025, including a 12–4 Week 15.

Sixty-three percent sounds much better than 52.8%. It isn't, and the reason is the benchmark. Against the spread, the bar is 50% by design. Straight-up, the bar is whatever picking the favorite every week gets you — and in the modern NFL that is roughly 66%. Measured against its own benchmark, our straight-up accuracy is below the trivial rule, and our own hold-out testing agrees: ~64% on unseen seasons, against that same ~66% baseline.

That is the honest read, and it is the one most services would bury: picking winners is an easier task, so a bigger number there means less, not more. We publish the straight-up projections because they are useful context on how the model sees a matchup — not because they beat the simplest possible strategy.

(We also tested an over/under objective. After months of feature engineering, totals graded out as near-perfectly calibrated — the hardest target we've pointed the machine at. The totals projections we publish carry our most conservative confidence tiers because of that research, and the full story is its own article.)

How the predictions are made

No gut, no lean, no Tuesday hunch:

  1. Every completed game since 2002 is decomposed into 360 computed trend features — rolling efficiency stats, situational splits, market-derived signals.
  2. A gradient-boosted model (with a logistic-regression cousin as a sanity check) is trained on 2003–2021, the feature search is scored on 2022–2024, and the promoted model is graded once on 2025–2026 — seasons the search never sees. Those three sets are kept separate precisely so the selection score can never be mistaken for the hold-out score again.
  3. Each week, the locked model re-trains its weights on the most recent window and scores the upcoming slate. Every prediction gets a 1–5 star confidence rating mapped from the model's own probability, so more stars mean the model is more confident — not that the outcome is more certain.
  4. Results are graded automatically against closing lines and published, hit or miss.

The model has no favorite teams, no narratives, and no memory of last week's embarrassment. That last property, it turns out, is the hardest one for humans to replicate.

Judge for yourself

The complete graded record — by season, by week, by confidence tier — is on the SpreadTrends Predictions page, updated every week of the season. We'd rather earn trust with a real 52.8%, and show the testing that produced it, than rent it with a fake 65%.

Record as of the end of the 2025 season: 821–733–43 against the spread lifetime; 148–134–3 in 2025; 85–50 straight-up since SU grading began (Week 11, 2025). Pushes excluded from percentages.