The Altitude Bet That Almost Convinced Us (And Two Other Near-Misses)

Analysis · by Rick's Picks Analytics

Every so often a hypothesis comes along that wants to be true. The story is clean, the recent numbers are loud, and you can already picture the write-up. This is a post about three of those -- an altitude angle, a special-teams signal, and a three-and-out stat -- and why all three are sitting in our "watch, don't act" folder instead of in the picks engine.

If you read this site regularly, you know our whole brand is that we publish the failures. This one is more interesting than a failure. These are near-misses: angles that look great in exactly the window where looking great means the least.

First, the ground rules

Quick refresher on how we evaluate any angle:

Near-miss #1: Thin air, thick sample problems

The hypothesis: when a team whose home stadium sits below 1,000 feet travels to play a host at 4,000+ feet of elevation, the visitor's conditioning suffers and the market underprices the home team. It's an old sharp-circuit favorite -- Wyoming at 7,220 feet, the service academies, the Mountain West in general.

Here's what the data says, and why it almost got us:

WindowRecordATS hit rate95% CIp vs break-even
Train (2015-2022)84-76 (n=160)52.5%44.8% - 60.1%0.52
Test (2023-2025)35-22 (n=57)61.4%48.4% - 72.9%0.109
2025 only12-5 (n=17)70.6%46.9% - 86.7%0.103

Look at that progression. 52.5%, then 61.4%, then 70.6%. Each time we re-ran the study with a new season of data, the recent number got better. When we first tested this angle, the test window (then 2023-2024) sat at 57.5% on 40 games. A year later the pooled test was 61.4% and the newest season alone was above 70%. If you only looked at the trend line, you'd swear the angle was heating up.

Now look at the other columns.

The train window -- 160 games across eight seasons, the biggest and most stable sample we have -- shows 52.5% with a p-value of 0.52. That is a coin flip. Not "marginal," not "suggestive." A literal coin flip. And the test window's confidence interval runs from 48.4% to 72.9%: it contains pure chance. The eye-popping 70.6% in 2025 comes from seventeen games. Seventeen. A 12-5 run happens to random strategies all the time; that's why its CI stretches down to 46.9%.

So the pattern is: zero signal in the large sample, loud signal in the small one. That is not what a real edge looks like. A real edge -- like the key-number margin distribution we wrote about previously -- shows up in train and persists in test. What altitude shows is the signature of noise that happens to be clustered in the recent window, which is precisely the window most tempting to overweight because it feels current.

There's also a market-logic reason to be skeptical: altitude is not a secret. Books have priced Laramie and Colorado Springs for decades. For this angle to be real, the market would have to be systematically mis-pricing one of the most publicly known situational factors in the sport. Possible? Sure. The way to bet on "possible"? Don't.

Near-miss #2: The special-teams convergence

This one is subtler and, honestly, the most interesting of the three.

We tested SP+ and FPI component ratings as spread predictors. The headline results were dead: overall SP+ edges hit 49.5% in both train and test; FPI's headline number was 49.7% train, 50.0% test. The market has these ratings fully absorbed -- no surprise, since they're free and famous.

But one component kept twitching in the test window, in two independent rating systems:

SignalTrainTest (2023-25)2025 only
SP+ special-teams gap50.05% (p=0.998)53.4% (p=0.167)56.0% (p=0.033 uncorrected -- fails Bonferroni)
FPI special-teams efficiency (threshold 10)49.3%54.3% (p=0.099)56.8% (n=361, p=0.052)

Two different providers, two different methodologies, both pointing the same direction in the same window: teams with a big special-teams rating advantage covered at 53-57% in the test era while showing nothing -- actually slightly worse than nothing -- in train.

The convergence matters. One test-only trend is noise. Two independent measurements of the same underlying thing trending together is at least a coherent story: maybe special teams is the component the market weights least, because it's the least glamorous and the hardest to see in a box score.

But run the checklist. Train support: none (50.05% and 49.3%). Multiple-comparison survival: the best p-value, 0.033 in the 2025-only SP+ slice, fails the Bonferroni bar. This is exactly the divergence pattern our rules exist to filter -- and our own history says the filter is right, because we've watched this movie before. Which brings us to...

Near-miss #3: The three-and-out head fake (and why it saved us)

Prior-season drive efficiency: touchdown rate, field-goal rate, score rate, plays per drive, three-and-out rate. The theory was that drive-level stats capture defensive quality that headline ratings smooth over.

Four of the five metrics were flat everywhere -- train hit rates between 49.3% and 50.4%, test essentially the same. Market priced. But three-and-out rate showed a wiggle: 50.1% in train, and in the 2023-2024 test window, 52.7% on 1,476 games. That was a big sample flirting with break-even, and we flagged it as worth watching.

Then 2025 happened. The 2025 season alone came in at 47.8%. The pooled 2023-2025 test now reads 51.1% (1,117-1,068 on 2,185 games) -- below break-even, headed back toward the train-window coin flip where it always lived.

This is the punchline of the whole post. Three-and-out rate gave us a genuine test-only trend on a large sample -- 1,476 games, far bigger than altitude's 57 -- and one more season of data refuted it anyway. If a test-only trend on 1,476 games can evaporate, what should we expect from one on 17?

We watched it without acting on it, and the discipline cost us nothing, because there was nothing there.

What would actually change our minds

"Watch, don't act" is not a euphemism for "quietly hope." Each of these has a concrete falsifiable condition:

The common thread: the only thing that promotes an angle is evidence in both windows. New test-era data can keep an angle on the watch list; it can never, by itself, put an angle in the engine. That's the rule that kept us off three-and-out before 2025 embarrassed it, and it's the rule keeping us off altitude now.

The bottom line

The altitude angle went 12-5 in 2025 and we didn't act on a single one of those games. That's not a missed opportunity -- that's the methodology working. Its eight-season train record is a coin flip, its test confidence interval spans pure chance, and we just watched a much larger test-only trend (three-and-out, 52.7% on 1,476 games) collapse to 47.8% the moment new data arrived. The special-teams convergence is genuinely intriguing and still fails every gate we have.

Recent, loud, and small beats old, quiet, and large in the human brain. It loses to it in the data, roughly every time. We'll re-run all three after 2026 and publish whatever we find -- including, as always, if the answer is "still nothing."

Rick's Picks publishes statistical analysis of college football for informational and entertainment purposes. Nothing here is betting advice. 21+. If gambling is affecting you or someone you know, call 1-800-GAMBLER.