We Fed Our Model a Season It Had Never Seen

Analysis · by Rick's Picks Analytics

There's a moment every quantitative model has to face eventually: fresh data it had no chance to memorize. For our engine, that moment was the completed 2025 college football season — 807 FBS games that didn't exist when any of our signals were promoted.

This week we re-ran every hypothesis file in the pipeline with 2025 added as holdout data. Train stays 2015-2022 (5,649 games); the test window extends to 2023-2025 (2,313 games, of which 807 are 2025). Break-even remains 52.38% at standard -110 pricing, and everything still has to clear Bonferroni-corrected significance. Same rules as always — the only thing that changed is that the model now has to answer for a season it had literally never seen.

Here's what happened.

The headline: all four standing signals survived

Going into this review we had exactly four live promotions: the Big Ten defensive totals adjustment, the cold-weather totals adjustment, CFB key-number confidence modifiers, and the QB metric (informational only). All four survived contact with 2025. Zero new signals were promoted. And two of the survivors got interesting in opposite directions — one strengthened dramatically, one had to be trimmed.

Cold games got colder: 53% retention became 94%

When we first promoted the cold-total signal (games under 40°F, non-dome, scoring runs below the market total), the numbers were honest but modest:

WindowEffect on totalSampleNotes
Train 2015-2022-2.73 ptsn=359p=0.007, Cohen's d=-0.15
Test 2023-2024 (at promotion)-1.46 ptsn=7753% retention — barely passed our 50% rule
Test 2023-2025 (now)-2.57 ptsn=11194% retention
2025 only-5.12 ptsn=34d=-0.31, the strongest single season yet

A signal that squeaked past promotion at 53% effect retention now retains 94% of its training effect — and 2025 alone was nearly double the training magnitude. That's the pattern you want from a real physical effect rather than a statistical mirage: more data, same direction, effect holds.

We're resizing the live coefficient from -1.4 to -2.0 — not to the -2.57 pooled estimate, and definitely not to the -5.12 from 2025. Why the restraint? The pooled test confidence interval is [-5.53, +0.39], which still crosses zero, and the 2025 slice is only 34 games. Sizing a coefficient to a 34-game season is exactly how you end up publishing a retraction next July.

Meanwhile, the neighboring weather signals stayed dead. Precipitation — which looked spectacular in train at -3.83 points (p < 0.0001) — sign-flipped again in 2025, coming in at +2.03. Wind over 15 mph still doesn't clear Bonferroni. Dome games still show nothing. Cold is the only weather signal that survives, and it keeps surviving.

The Big Ten under got stronger

The one legacy engine coefficient that survived last year's audit — Big Ten games running under the posted total — didn't just hold in 2025. It nearly doubled out of sample:

WindowEffectSignificance
Train 2015-2022-2.68 ptsp=6e-06, d=-0.15
Test 2023-2025-4.27 ptsp=1.3e-06, d=-0.26, n=401
2025 only-2.90 ptsp=0.052, n=136

The test effect is stronger than the training effect, and the 2025-only slice lands at 108% of the training magnitude with the same sign. Here's the discipline part: we are not raising the coefficient. The live value stays at the train-derived -2.68. Chasing the test number would be fitting to holdout — the exact sin the holdout exists to catch. If anything, the live value is now conservative, and that's fine.

For completeness: all nine legacy coefficients we zeroed out last year (SEC advantage, dome bonus, rain penalty, home-favorite fade, and the rest) stayed dead in 2025. The best fade among them hit at 48.2-54.5% — nothing near significance.

The part where we demote our own findings: key numbers trimmed

This is the entry we most want you to read, because it's our own published work getting cut down.

In an earlier post we walked through CFB key numbers — the margins where games actually finish — and our engine used a set of seven: 3, 7, 14, 17, 21, 24, 28. With 2025 in the data, we went back and asked a harder question: which of these are Bonferroni-robust spikes above the empirical baseline, and which are peak-picking artifacts?

MarginTrain PTest 2023-25 P2025-only PVerdict
39.6%10.7%11.4%Keep (p_bonf=2e-95, h=0.26)
78.5%8.7%8.3%Keep (p_bonf=2e-68, h=0.22)
144.4%4.2%3.8%Keep at half weight (p_bonf=0.0066, h=0.048 — tiny)
173.4%4.2%4.8%Drop — never passed Bonferroni in train (p_bonf=1.0)
213.8%3.8%2.6%Drop — occurs below the uniform baseline in 2025 (ratio 0.75)
243.0%2.9%2.9%Drop — below uniform in 2025 (ratio 0.83)
2822.0%19.3%19.3%Reclassified — this is the clipped "28 or more" blowout bucket, not a key number

Three and seven are bedrock — enormous effect sizes, replicated in every window, top-10 margin overlap of 90% between train and test and between train and 2025-only. But 17, 21, and 24 were riding along on the strength of their neighbors. In train they never cleared Bonferroni. In 2025, margins of 21 and 24 actually occurred less often than a uniform distribution would predict. They were artifacts of eyeballing peaks in a histogram, and the honest move is to say so and cut them.

The ±5 / ±3 confidence-modifier math is unchanged; the set of numbers it keys on shrinks from seven to effectively {3, 7, 14-at-half-weight}. And 28 keeps its one legitimate job: it remains the only margin where a half-point buy at -120 carries positive expected value, because it's really a blowout-tail bucket.

What 2025 killed, and what it almost resurrected

Every previously rejected signal stayed rejected — 2025 was brutal to the "almosts" from prior sweeps. The bye-week rest angle, which was our closest near-miss at 53.0% train, hit 46.2% in 2025. The pace-based OVER angle's 53.0% test trend collapsed to 47.7%. The three-and-out defensive proxy fell from a hopeful 52.7% to 47.8%. The moneyline value model that showed +14.7% train ROI came in at -6.7% in 2025 — the classic overfit signature completing its arc.

Two verdicts changed, neither into a promotion. The home-QB-out effect improved its retention from 41% to 77% (train -1.98, test -1.53, 2025 -3.07 on 54 games) and moves from rejected to informational — but it's blocked from the engine because the away-side effect is exactly 0.00 in train, which smells like confounding, and no ATS backtest exists to score it against break-even yet. And the altitude angle — home teams at 4,000+ feet against lowland visitors — hit 70.6% in 2025. On 17 games. With zero training-set support (train p=0.52). It goes on the watch list, not into the engine. A 17-game sample proving a theory the training data rejected is a story, not a signal.

The bottom line

We fed the model 807 games it had never seen. The four signals we'd staked our name on all survived, one of them (cold totals) strengthened from marginal to solid, and one of them (key numbers) forced us to publicly trim our own published finding because three of our seven "key numbers" turned out to be artifacts. Zero of the twenty-plus rejected hypotheses earned their way back in, and several near-misses were finished off for good.

That's the whole point of holdout validation: it doesn't care what you published. The market priced almost everything we tested, the handful of real effects are small and physical (cold weather, Big Ten scoring environments, the geometry of football margins), and the discipline of demoting your own findings is the only thing separating analysis from tout-speak. Next checkpoint: September 2026, when live picks start generating fresh out-of-sample data all over again.

Rick's Picks publishes statistical analysis of college football for informational and entertainment purposes. Nothing here is betting advice. 21+. If gambling is affecting you or someone you know, call 1-800-GAMBLER.