The Promotion Gauntlet: How a Signal Earns Its Way Into Our Picks
Analysis · by Rick's Picks Analytics
Every pick you see on this site is powered by signals that survived a filter designed to kill them. We call it the promotion gauntlet, and the numbers tell you how brutal it is: in one sweep of nine team-quality signals (SP+, FPI, recruiting composites, advanced metrics, and more), zero of nine were promoted. In a follow-up sweep of seven classic "sharp angles" -- altitude, revenge spots, line movement across key numbers -- zero of seven. As of today, out of roughly thirty hypothesis files we've built and tested, exactly four signals are live in the engine.
That failure rate isn't embarrassing. It's the product. Here's how the sausage is made.
Stage 1: The train/test wall
Every hypothesis gets fit on games from 2015-2022 (the train set, 5,649 FBS games) and then scored on games from 2023-2025 (the test set, 2,313 games) that it has never seen. The split is temporal, not random, because random splits leak: if your model trains on October 2023 and tests on September 2023, it's quietly using the future to predict the past.
We learned this the expensive way. Our first SP+ run showed a 60-62% against-the-spread hit rate at high thresholds -- spectacular, if true. It wasn't. We were using same-season SP+ ratings as the predictor, which means the rating already "knew" how the season went. After shifting to prior-year ratings, the edge vanished to 49.5%. We documented that in the ledger specifically so nobody ever re-reads the old numbers and believes them.
Stage 2: Significance, corrected
A p-value under 0.05 sounds impressive until you remember we test dozens of angles on overlapping data. Run twenty tests and one will look "significant" by pure chance. So we apply a Bonferroni correction: when a file tests multiple related hypotheses, the significance bar tightens proportionally.
This kills real-looking candidates. Wind over 15 mph suppressing totals? Train p = 0.13 -- dead on arrival under our corrected threshold. A turnover-luck fade that hit 55.8% in train never actually cleared its corrected bar (p = 0.074 against a threshold of 0.025), and sure enough it hit 45.0% in the 2025 holdout.
Stage 3: Effect size, because significance alone is noise
With 5,649 training games, tiny effects can be statistically significant and practically worthless. So every test reports Cohen's d (or an equivalent effect size) next to its p-value. Our cold-weather signal carries d = -0.15 in train -- small but real, and worth about 2.7 points on a total. Margin-of-3 games pass Bonferroni at p = 2e-95 with effect size h = 0.26; margin-of-14 games also pass Bonferroni (p = 0.0066) but with h = 0.048, which is why 14 gets reduced weight in the engine while 17, 21, and 24 -- which never passed Bonferroni at all -- got dropped entirely.
Stage 4: Sign retention and the 50% rule
Here's where most survivors die. A signal that passes significance in train must show up in test with the same sign and at least 50% of its train magnitude.
Sign flips are automatic kills, no appeal. If rain suppressed totals by 3.8 points in train and inflated them in test, one of those windows is lying, and we have no way to know which.
The 50% rule catches slower deaths. Our returning-production angle showed low-returning-production home teams covering just 45.5% in train -- a fade-home edge of +6.9% above break-even. In test the edge shrank to +1.5%. Same sign, but 22% retention. Killed. Home teams with their QB out underperformed the market by 1.98 points in train (p = 0.006, n = 477) but only 0.82 in the 2023-24 test window -- 41% retention. Killed.
Stage 5: The market backtest
Even a stable, significant effect has to beat the actual betting market. At standard -110 pricing, break-even is 52.38%. Our moneyline-value signal is the cautionary tale: train ROI of +14.7% on 695 games looked like a career-maker. Test ROI: +1.3% on 700 games, and after adding the 2025 season, -1.1%. Classic overfit signature -- the train window memorized noise and called it edge.
A tale of two weather signals
Both of these started in the same hypothesis file, walters_weather_adjustments.py, with the same prior: bad weather suppresses scoring. The gauntlet split them.
| Stage | cold_total (temp < 40°F, non-dome) | precip_total (precipitation) |
|---|---|---|
| Train effect on totals | -2.73 pts (p = 0.007, d = -0.15, n = 359) | -3.83 pts (p < 0.0001) |
| Significance under Bonferroni | Passed | Passed -- looked stronger than cold |
| Test (2023-24) | -1.46 pts on 77 games (53% retention, same sign) | +1.64 pts -- sign FLIPPED |
| 2025 holdout | -2.57 pts pooled test (94% retention); 2025 alone: -5.12 (d = -0.31) | +2.03 in 2025 -- flipped again |
| Verdict | Promoted. Live as a -1.4 point totals adjustment, now being resized to -2.0 | Killed. Zero contribution to any pick |
Read that middle row carefully. On the training data alone, precipitation was the better signal -- bigger effect, smaller p-value. If we shipped signals based on train performance, rain would be moving our totals projections right now, in the wrong direction. The only thing that saved us was the wall between train and test.
And note what we did not do when cold_total came back at -5.12 points in the 2025 season: we did not resize the coefficient to -5. That's a 34-game sample. The recommendation on the ledger is -2.0 -- conservative, because the test confidence interval [-5.53, +0.39] still crosses zero. Chasing your best holdout year is just overfitting with extra steps.
Stage 6: The ledger
Every decision -- promote, reject, defer -- goes into a permanent integration log with the numbers at decision time. When a signal goes live, we record its test-set performance at promotion, then re-score it against every new season. The 2025 season was the first time our promoted signals faced a year that didn't exist when they were promoted. All four survived; cold_total actually strengthened (retention rose from 53% to 94%).
The ledger also forces us to keep score on our failures. It's where we admit that the 81.8% train hit rate on lines crossing key number 3 (n = 11) collapsed to 44.7% in test -- a fluke, confirmed. It's where near-misses like the altitude angle (61.4% test on n = 57, but zero train support) sit on a watch list instead of sneaking into the engine. And it's where we audited our own legacy engine and zeroed out nine of ten hardcoded coefficients -- including a "SEC advantage" that covered 47.1% in test and a "dome advantage" whose train effect was -0.05 points. We published the demolition of our own old code.
Why so few survivors?
Because the closing line is genuinely hard to beat. Every market-implied team-quality signal we tested -- SP+, FPI, talent composites, recruiting, twelve flavors of advanced metrics -- landed between 49% and 51% against the spread once look-ahead bias was removed. The market has already absorbed them. The four survivors (key numbers, cold totals, a Big Ten defensive totals effect, and an informational QB metric) share a trait: they're structural properties of how the sport is played or narrow situational effects, not "this team is better than the market thinks."
The bottom line
The gauntlet is: temporal train/test split, Bonferroni-corrected significance, effect sizes alongside p-values, sign retention with a 50% magnitude floor, a backtest against the 52.4% break-even, and a permanent ledger of every verdict. Most signals die, and we publish the obituaries. When a projection on this site carries a cold-weather adjustment or a key-number confidence modifier, it's because that signal beat data it never saw -- twice. That's not a promise of profit; it's a promise of process. In this business, the process is the only thing you can actually control.
Rick's Picks publishes statistical analysis of college football for informational and entertainment purposes. Nothing here is betting advice. 21+. If gambling is affecting you or someone you know, call 1-800-GAMBLER.