NFL Luck: The Regression Claim We Measured and Won't Make

The 2024 Kansas City Chiefs went 15-2. Their efficiency — net expected points added per play — was the profile of a 9-win team. A year later they went 6-11.

The 2022 Minnesota Vikings went 13-4 while being out-played per snap across the season: net EPA of −0.134, the profile of a 7-win team. A year later they went 7-10.

Those two seasons are why every offseason produces the same list: the lucky teams due to come back to earth, and the unlucky teams due to bounce back. It is one of the most repeated forecasts in football analysis, and it is presented as symmetric — the lucky fall, the unlucky rise.

Our Rankings page now carries a luck figure for every team, and before we added it we measured whether the list means what it says. It half does. This article is about the half that doesn't, and why we will not publish it.

⚠️ Everything in this article is straight-up — actual wins and losses, the point spread ignored — except one section near the end that is labelled as being about the spread. "Record", "win%" and "wins" all mean games won, not games covered.

What "luck" means here

A win-loss record is a noisy measurement of how good a team is. Seventeen games is not many, a handful of them are decided by a field goal, and fumbles bounce where they bounce.

Expected points added (EPA) is a steadier measurement. On every snap it asks how much the play changed the offense's expected points given down, distance and field position, and it does not care who won the game. Across 2002–2025, a team's net EPA and its straight-up win percentage move together very closely in the same season (r = +0.845), so EPA gives a decent estimate of the record a team "should" have had.

Luck, as we define it, is the distance between the two: actual wins minus the wins that team's EPA implies. Positive means the record ran ahead of the play. Negative means it ran behind.

How big the gap actually is

Before asking whether the gap predicts anything, it helps to know how large it gets. Across 768 team-seasons:

Distance between a team's straight-up record and its EPA-implied record Share of team-seasons
Under half a win 26%
One win or more 55%
Two wins or more 22%
Three wins or more 7%

The typical team finishes about 1.1 wins away from its EPA-implied record, in one direction or the other. The extremes run past five:

Season Team Straight-up record EPA-implied wins Gap
2022 Vikings 13-4 6.9 +6.1
2024 Chiefs 15-2 9.3 +5.7
2012 Colts 11-5 6.0 +5.0
2011 Packers 15-1 10.3 +4.7
2017 Browns 0-16 5.5 −5.5
2025 Chiefs 6-11 10.2 −4.2
2019 Cowboys 8-8 12.1 −4.2
2016 Jaguars 3-13 7.1 −4.1

Hold on to that 1.1 wins figure. A different, smaller number is coming, and the two are easy to confuse.

The first-order effect: records just regress

Here is the part that makes every regression list look right.

Bad teams get better the next year, and good teams get worse, whatever their efficiency said. Split the bad teams by luck and both halves rebound; split the good teams and both halves decline:

Teams that won… Their luck said… Change in straight-up win% next season In 17-game wins
6 or fewer they were unlucky +.181 +3.1
6 or fewer they were simply bad +.133 +2.3
10 or more the record outran the play −.164 −2.8
10 or more the play backed the record −.109 −1.9

A 3-14 team finds two or three more wins without anyone having to explain why. That is ordinary regression to the mean, and it is enormous next to anything luck adds.

It is also why a raw luck leaderboard is misleading on its face. Extreme records attract extreme gaps: 10-plus win teams average +1.0 wins of luck and 6-or-fewer teams average −1.0, simply because a 13-win season is more likely to have had close games break its way than a 13-win-quality team is to exist. Rank teams by luck and the "lucky" side fills with good teams that were going to decline anyway. Their decline then looks like the luck catching up.

So the only fair test holds the record roughly constant and asks whether the luck figure adds anything on top of that regression.

The second-order effect, at the bottom: about 0.6 wins

Among teams that won 6 or fewer, the half whose EPA said they deserved better finished the following season at .449; the half whose EPA agreed they were bad finished at .416. That is about 0.6 wins over a 17-game season.

We checked it a second, stricter way — comparing teams only against others with the identical W-L-T record, so the two halves start from exactly the same place. The gap held: .454 against .417, or 0.63 wins.

Two methods, the same answer. That is a real effect, and it is worth being precise about what it is:

⚠️ 0.6 wins is not the size of a team's luck. It is the next-season difference between two groups of teams that were already bad, split by this measure. A team whose record trailed its EPA by four wins should not be expected to gain four wins back, or 0.6 wins back, from luck alone. Its rebound is mostly the ordinary +2.3 to +3.1 every bad team gets, with a fraction of a win on top.

The second-order effect, at the top: nothing we can stand behind

Now the good teams. Among teams that won 10 or more, the half whose record outran its EPA finished the following season at .573. The half whose EPA backed the record finished at… .573. The same number, to three decimals, across 254 team-seasons.

The lucky half declined further, but only because it started higher. Where it landed was identical.

The stricter identical-record test disagreed: it put the lucky half 0.49 wins below the deserving half the next year. So one method says zero and the other says half a win.

That disagreement is the result. An effect that appears or vanishes depending on how you cut the sample is inside the noise of the cutting. It is not something to put next to a team's name.

Why the list is right in one direction and not the other

Put the two halves together and the popular regression list is:

  • Defensible for unlucky bad teams. Their rebound is mostly ordinary regression, but there is a small, stable extra amount that two different methods agree on.
  • Not defensible for lucky good teams. Their decline is real but is the same decline every good team suffers. We cannot show that the luck figure tells you which good teams will fall further.

The Chiefs and Vikings collapses at the top of this article are the cases everyone remembers. The 2012 Colts are the case nobody does: 11-5 on an EPA profile worth six wins, then 11-5 again the following year, once more well ahead of their efficiency. Across the full cohort, the 10-win teams that "got lucky" and the ones that didn't ended up in the same place.

The popular version of the list treats both directions the same. The data does not.

A check from the other direction: the point spread

There is a second way to define luck that a site built on spread data might reach for first: how far a team finished from the closing point spread, game by game. A team that kept finishing ahead of the market's number had, in some sense, outrun expectations.

Split broadly, that version looked like it worked. But the halves it produced started 7.1 win-percentage points apart in the bad-team group — it was mostly sorting by record, which brings back the regression-to-the-mean problem. Held to identical W-L-T records, it collapsed to −0.02 wins for good teams and +0.08 wins for bad teams. Essentially zero.

That is not a failure; it is a consistency check that passed. Finishing ahead of the market's number is a real thing to measure, but if it did carry into the following season, the market would be leaving something predictable on the table year after year. It does not, which is what you would expect from a number that already prices in everything the public knows. So the luck figure we use is the efficiency version, and it is about straight-up records only.

What the Luck column on our Rankings page says

That is why our Luck column describes the season rather than forecasting the next one. It sits on the EPA Ratings tab and shows, in wins, how far each team's straight-up record has run from the record its net EPA usually produces: +1.6 means about one and a half more wins than the play supports, −2.3 about two fewer. Regular-season games only, and it stays blank until teams have played four games, because before then a single result can move it by a whole win.

  • A positive value means "record ahead of efficiency." It does not mean "decline candidate", because we measured that claim and could not support it.
  • A negative value means "record behind efficiency." For a team that is also losing, history adds about 0.6 wins of next-season rebound on top of the ordinary bounce-back. That is context, not a forecast for any one team.

The column has no colours and no 1–32 rank, on purpose. Most teams sit within a win of zero, where a rank would imply an order the numbers do not support, and colouring one side green and the other red would make exactly the symmetric claim this article argues against.

The honest caveats

  • The sample is not 736 independent trials. It is 32 franchises observed repeatedly, so the same organisations, front offices and quarterbacks show up across many pairs.
  • Matching on record controls record only. It does not hold roster, quarterback, coaching or schedule constant, and any of those can swamp half a win.
  • The EPA-to-wins conversion is fixed, fitted once on 2002–2025 (win% = 0.5088 + 0.7769 × net EPA). We did not search for a better formula, and a different fit would move individual teams.
  • Regular season only, 2002–2025, ties counted as half a win. 2026 is excluded from every figure in this article; the Rankings column covers it as the season is played.

See it yourself

Every team's Luck figure, for this season and every season back to 2002, is on the SpreadTrends Rankings page under the EPA Ratings tab, next to the four-way EPA splits it is built from — offensive and defensive, pass and rush. Sort the column to find this year's largest gaps in either direction, and read them with the asymmetry above in mind. The formula and how to read it are in our FAQ. For the team-by-team view of the same gap from the 2025 season, see The EPA Mirage.

Data: 768 team-seasons 2002–2025 from the SpreadTrends Rankings and Games tables; 736 year-over-year pairs. Expected wins from a same-season fit of straight-up win percentage on net EPA (r = .845). Regular season only.