Why We Don't Have a Strong Take on Every Game
Every NFL week has around sixteen games, and by Thursday the content industry has produced sixteen confident opinions about each of them — a winner, a score, a reason, delivered with the same certainty whether it's a coin-flip divisional game or a three-touchdown mismatch.
Our model looks at the same sixteen games and, most weeks, says something far less marketable: on most of them, it isn't sure.
That's not a flaw we tolerate. It's the most honest output the system produces, and the reasoning behind it is the single most transferable idea we can offer anyone who follows football through numbers.
Certainty is a distribution, not a personality
When our model scores a game, it doesn't produce a take — it produces a probability. And here is the uncomfortable truth about NFL forecasting that sixteen-confident-takes content hides: most of those probabilities land near 50/50. The league is engineered for parity, the schedule is engineered for drama, and the market number each game is measured against has already absorbed most of what's knowable. After all of that, genuine forecastable daylight is scarce — and it is unevenly distributed across the slate.
So we measure it, on a scale that does not grade on a curve. Every projection carries a confidence reading from 1 to 5, mapped from one thing only: how far the model's probability landed from a coin flip. The cutoffs are fixed in advance and never move:
| Reading | Model's projected probability | What we call it |
|---|---|---|
| 5 | 78.3% or better | High-confidence |
| 4 | 69.9% – 78.3% | High-confidence |
| 3 | 62.8% – 69.9% | Solid |
| 2 | 56.3% – 62.8% | Solid |
| 1 | under 56.3% | Lean |
That fixity is the whole design. A 4 in Week 2 and a 4 in Week 15 are the same claim about the same distance from 50/50, so a reading means one thing across a season and across seasons — you can compare them, and you can hold us to them.
The obvious alternative is to rank each week's board and hand out the readings by position: top quarter takes the highest, next quarter the one below, the rest at the bottom. That's tidier to present, and we deliberately don't do it, because a scale built that way guarantees a confident-looking projection every single week whether or not the model found anything. It can never come back empty at the top, which means it isn't measuring conviction — it's distributing it on a schedule.
Ours can come back empty at the top, and does. On the 2026 opening slate, the model's sixteen against-the-spread projections came back as nine readings of 1, six of 2, and a single 3 — nothing at 4 or 5 anywhere on the board. Fifteen of the sixteen biggest-audience games of the year sat at Lean or Solid, and the model's firmest spread opinion all week was a 3 of 5. A relative scale would have printed four confident-looking calls that week regardless. Ours printed none, because there were none.
A reading of 1 is the model saying: I produced a lean because you asked, but the number I'm projecting against is close enough to my own that I don't know, and neither does anyone else. On most weeks that describes a large share of the board. We consider that a feature. It might be the most rigorous sentence on the site.
What one confident week actually proves
The math of why uniform confidence is a tell — worked through honestly — is worth internalizing.
A forecaster with zero skill picking sixteen games gets eight right on an average week, and every few weeks — by pure arithmetic — lands eleven or twelve. That's not a hot streak; that's what a coin does. Our own graded history makes the point against ourselves: the same model, the same process, produced a 13–1 week in October 2025 and a 32–45 stretch across five weeks in December. If we'd shown you only the first, you'd think we were oracles. Only the full record, streaks and slumps both, tells you what the system actually knows — which is why we publish it.
Now apply that lens to the sixteen-takes industry. A source with a confident opinion on every game is guaranteed regular 11–5 weeks to screenshot, no skill required, and no memory required either — because takes without a graded ledger reset to zero every Thursday. Uniform confidence isn't analytical strength. It's what the absence of self-measurement looks like.
Selectivity is where skill actually lives
Here's the deeper principle, and it applies to any forecasting domain, not just football: a forecaster's skill isn't how often they're right — it's whether their confidence means anything. A weather service that says "70% chance of rain" and is right 70% of those times is calibrated, and calibration is the entire difference between a forecast and a vibe.
That's the standard we hold the scale to, and an absolute scale is what makes the standard testable. Games the model read at 4 or 5 should, over seasons, grade out meaningfully better than the ones it read at 1 — and because every projection is stored and displayed with the reading it was given, anyone can check whether they do. A relative scale would quietly frustrate that check: its top bucket is repopulated every week by construction, so it is never empty and never comparable across weeks. Fixed cutoffs put our calibration on the record, permanently. A take can't be calibrated. It isn't even wrong; it's just gone by Monday.
There's also a quieter benefit, one every serious forecaster learns eventually: the discipline of saying "I don't know" is what keeps the "I know" honest. A system forced to publish its own certainty on a scale it cannot curve can't hide its weak opinions inside its strong ones. Neither can a person, once they start keeping score.
How to use this — on our site or anywhere
- On SpreadTrends: the confidence meter on each projection is the model's own certainty, on a scale whose readings mean the same thing in every week of every season. A 4 or a 5 is where the numbers found something; a 1 is the model's honest shrug. Both are information — and a week with nothing above a 3 is information too, which is exactly why we let that happen rather than curving it away.
- Everywhere else: count the takes. When you find a source with sixteen confident opinions a week, ask the only diagnostic question that matters — where's the graded ledger, and where's the reading that says "unsure"? No ledger and no "unsure" level means you're reading entertainment. Which is fine — football is entertainment — as long as nobody's calling it a forecast.
The sharpest people in any prediction field share one habit: they are violently selective about what they claim to know. The NFL slate gives you sixteen chances a week to have an opinion. The numbers, most weeks, are firm about far fewer — and some weeks, about none at all.
Every projection our model makes — with its confidence reading and its graded result — is on the SpreadTrends Predictions page.