The math.
The evidence.
The payout conversion is exact for a supplied distribution. The distribution estimates how often a game finishes on each margin. These are different kinds of accuracy.
← Back to the NFL calculatorThe calculator treats your visible original spread as the main game spread. It has no hidden overrides or custom probability settings. Original odds set the payout whose modeled return is preserved; they do not identify the true probability or move the main game spread. The ML row uses a two-way full-game moneyline, including overtime, with a tied game refunded. Its price is calculated from the same score-margin distribution.
1. Define equivalent precisely
We preserve expected net return for $1 risked. Let W, P and L be the probabilities of a win, push and loss; D is decimal odds. A push returns the stake.
= W × D + P − 1
For reference probabilities W₀, P₀ and decimal odds D₀, the target price with target probabilities W₁, P₁ is:
Model fair decimal odds = (1 − P₁) / W₁
The first preserves the reference bet's modeled edge or cost. The second is break-even under the distribution. They coincide only when the reference bet is model-fair. Invalid nonpositive payouts are unavailable rather than forced into valid odds.
2. Estimate integer score margins
We use public nflverse game data: completed regular-season and playoff games from 2015–2025. Scores, unique game IDs and spread signs are checked against the official field dictionary. The canceled 2022 game and unfinished games are excluded. The data snapshot contains 3,028 eligible games; 2026 is excluded because the season is incomplete.
Each past game is oriented to its market favorite. Games receive Gaussian weights according to distance between their main spread magnitude and your game's main spread magnitude. The selected bandwidth is 2.5 points. Each actual integer margin retains its own probability mass; there is no smoothing across margins or clipping at ±24.5.
At a zero main spread, both orientations of each game are averaged. Historical pick'em games are also averaged within the same game. They never count as two independent games. Positive team spreads mirror the favorite's margins.
The calculator assumes a balanced main spread: 50% cover probability conditional on no push, equivalent to a proportional de-vig anchor for −110 / −110. Win and loss groups rescale to that anchor while the estimated main-line push probability is retained. Your original price then determines the return preserved by the conversion. A price such as −126 can reflect unequal market probabilities, extra vig, or a different underlying main spread; a single quote cannot identify which. The simple calculator does not infer that information. Use the game's main spread as the original spread; an already-alternate original line can give a misleading context.
3. Choose settings before the test
| Phase | Seasons | Games |
|---|---|---|
| Initial fit | 2015–2020 | 1,604 |
| Parameter selection | 2021–2022 | 569 |
| Untouched evaluation | 2023–2025 | 855 |
Sixteen candidates combined spread bandwidths 1.5, 2.5, 4 and 6 with total bandwidths 4, 8, 16 or no total conditioning. The predefined score averaged the three-outcome Brier score at nine offsets from each main spread: −6, −3, −1, −0.5, 0, +0.5, +1, +3 and +6. Lower scores are better.
The selected model used spread bandwidth 2.5 and no total conditioning. Total conditioning made no improvement in this validation, although the differences were small. A normal-margin baseline with width 12 was independently selected from widths 12–15 using the same validation period. Both models use the same market anchor. Parameters were saved before examining held-out scores.
For the held-out test, historical probabilities used only games through 2022. The deployed calculator subsequently refits the selected model on all 2015–2025 games. Its training sample is therefore larger than the model evaluated in the holdout.
4. Held-out results
| Measure | Result |
|---|---|
| Discrete-margin model Brier | 0.48949 |
| Normal baseline Brier | 0.49084 |
| 95% interval, model minus baseline | −0.00297 to +0.00027 |
| Whole-line push prediction / observation | 3.62% / 3.84% |
The observed improvement is small, and the interval includes zero. This test does not establish superiority over the normal baseline. The interval uses 2,000 paired resamples of complete games. Multiple spreads from the same game are correlated; 7,695 evaluated spreads do not mean 7,695 independent games.
Win-probability calibration
| Predicted band | Average prediction | Observed wins | Spread evaluations |
|---|---|---|---|
| 20–40% | 32.65% | 34.12% | 1,574 |
| 40–60% | 49.64% | 51.14% | 4,941 |
| 60–80% | 66.07% | 68.39% | 1,180 |
Observed favorite win rates were about 1.5–2.3 percentage points higher than the predictions in these bands. Empty bands below 20% or above 80% supply no calibration evidence. These descriptive bins were not used to retune the model after the test.
Key-number calibration
| Outcome | Average prediction | Observed frequency |
|---|---|---|
| Favorite wins by 3 | 8.30% | 8.19% |
| Favorite wins by 7 | 5.38% | 5.73% |
Each row evaluates 855 games. These checks assess pooled frequencies, not exact probabilities for any single matchup or main spread. Full per-season and spread-band results are in the numerical report.
5. Sampling and supported ranges
Effective sample size is (sum of weights)² / sum of squared weights, counted at the game level before market anchoring. It describes local historical support. It is not a guarantee that the model is correctly specified.
Original spreads are restricted to ±20. Estimates with fewer than 100 effective games, fewer than 2% target wins, or impossible positive payouts are unavailable. The earlier evaluation covered offsets within six points of the main spread. Larger moves and probabilities outside the observed calibration bands are exploratory extrapolations, even if the slider displays a price. Advanced controls and custom probabilities have been removed. Only the visible original spread, original odds and target line are saved in this browser.
6. Billy Walters reference check
We checked the supplied book tables on pages 266 and 269 of Billy Walters' Gambler. Page 266 gives rounded half-point dollar adjustments. At 3, the onto-push and off-push figures differ. Page 269 pairs a posted spread of 3 at −115 with 3.5 at +105 in an implicit-spread example. The original additive calculator returned +107 for that pair, so it was not an exact reproduction.
Fixed cents also fail to preserve expected return at arbitrary starting odds. For example, moving a win into a push removes the profit payout, whereas moving a push into a loss loses the stake. The adjustment therefore depends on both the reference payout and the margin distribution. This version uses the explicit expected-return equations above. It is an independent historical model, not a certified implementation of the entire book.
7. Full-game total follow-up
On October 4, 2026, we compared spread-only estimates with 16 total-aware alternatives. Spread bandwidth stayed fixed at 2.5; total bandwidths were 4, 8, 12 and 16 points, with 25%, 50%, 75% or 100% weight on the total-conditioned distribution. Settings were selected on 569 games in 2021–2022 after fitting 2015–2020. The score weighted nearby-line prediction 40%, moneyline 30% and key-number crossings 30%.
The nearby checks used ten offsets within ±4. Key checks used margins 2.5, 3, 3.5, 6.5, 7 and 7.5 only when within four points of the main spread. Moneyline used an exact zero threshold, with ties as a separate outcome. Both candidates retained the same two-sided market anchor for comparison.
The selected total candidate used bandwidth 8 and full total conditioning. It was then fitted through 2022 and evaluated on 855 games in 2023–2025. These seasons were previously examined in the first audit, so this is a retrospective comparison, not a fresh unseen holdout. No settings were retuned after this follow-up evaluation.
| Prediction | Spread only | With total | 95% difference interval |
|---|---|---|---|
| Nearby spreads | 0.496804 | 0.496756 | −0.000152 to +0.000057 |
| Moneyline | 0.421427 | 0.421412 | −0.000166 to +0.000137 |
| Key crossings | 0.506536 | 0.506520 | −0.000131 to +0.000098 |
Lower three-outcome Brier scores are better. Differences are total minus spread-only, using 5,000 paired whole-game bootstrap resamples. Key crossings included 800 qualifying games. All intervals include zero. The total candidate slightly improved all pooled scores but worsened the weighted score in games totaled at 47 or higher, and in the 2024 season.
The rule written before evaluation required a better weighted score, no worse pooled primary endpoint, and a 95% improvement on at least one endpoint. The candidate failed the last requirement, so the live calculator retains spread-only estimates and has no total input. This does not establish that totals have no effect, or that other models cannot capture it. A projected total is a forecast, not the game's realized total. The backtest uses historical market totals.
8. What remains unproven
Historical alternate-line quotes were not available in this dataset. We did not simulate returns against invented prices. Neither available-price accuracy nor betting profitability has been backtested. A future quote test would require timestamped main and alternate spreads, prices on both sides, consistent settlement rules, and actual outcomes, with parameters frozen before the tested period.
The main spread is archived market information. This is not a test of an opening-line strategy or live-game prices. Team strength, injury, weather, total and future NFL rules can change score-margin probabilities. Games also share teams and seasons; game resampling does not account for that dependence. Proportional de-vig and historical comparability may also be wrong. A narrow sampling range excludes those sources of uncertainty.
Source snapshot SHA-256: 88a64e3bacef4aba63da6ed7c9ea243a2e69b0b12e6d88d3ae01452f5b49cdc9. Retrieved October 3, 2026, New York time. Public numerical data and results are included with this site; private chat photos are not.