We Backtested Our Own Picks Engine. It Lost.

Analysis · by Rick's Picks Analytics

Most analytics sites will show you their winners. We're going to show you our backtest — the whole thing, including the part where our own engine loses to the closing line for eight straight seasons.

Rick's Picks is the prediction engine behind this site. It scores every FBS game on ranking gaps, conference strength, travel, rest, weather, key numbers, and a stack of situational factors, then assigns a confidence score and, when the edge is big enough, a pick against the spread. This offseason we did the thing every projection shop should do and most never publish: we replayed the live engine against 11 seasons of closing spreads and final scores, exactly as it would have run at kickoff, and graded every pick.

Here's what came back.

The headline numbers

We loaded 11,223 completed games from 2015-2025 with a closing spread and a final score. The engine fired a spread pick on 5,490 of them. At standard -110 pricing, the break-even hit rate is 52.38%. Following our usual methodology, we split by era: 2015-2022 is the period the engine's constants were built against, and 2023-2025 is genuinely out-of-sample.

EraPicks (n)Record (W-L-P)Hit rateROI at -110p vs break-even
2015-2022 (in-sample)3,8331874-1959-7848.9%-6.7%0.9999
2023-2024 (test)1,047517-530-1949.4%-5.7%0.976
2025 (out-of-sample)510274-236-353.7%+2.6%0.287

Read that first row again. On the very data the engine's rules were tuned around, it hit 48.9% against the spread — worse than a coin flip, and nowhere near the 52.4% you need just to cover the vig. That's a -6.7% ROI in-sample. In-sample is where models are supposed to look too good. Ours managed to lose money on its own home turf.

The 2023-2024 test window told the same story: 49.4%, -5.7% ROI. The 95% Wilson interval on that era runs from 46.4% to 52.4% — the top of the interval barely grazes break-even.

But wait — didn't it win in 2025?

Yes, and this is exactly where a less careful site would start the victory lap. In 2025 the engine went 274-236-3, a 53.7% hit rate and +2.6% ROI on 510 picks. Above break-even! Chart it in green! Crop out everything before September 2025!

Here's why we won't. The binomial test against the 52.4% break-even rate gives p = 0.29. In plain English: if the engine's true skill were exactly break-even — no edge at all — you'd see a single-season run this good or better about 29% of the time by pure chance. One good season on 510 picks is not evidence of an edge. It's a coin that happened to land heads a few extra times. The Wilson interval on 2025 runs from 49.4% to 58.0%, and 49.4% is exactly what the engine did the two years before.

There's also a detail inside the 2025 numbers that should make anyone skeptical: the engine's lowest-confidence picks (confidence under 60) hit 61.8% (n=55), while its highest-confidence picks (70+) hit 45.6% (n=57). If the confidence score meant what it claims to mean, that ordering should be reversed. When your model's certainty is inversely related to its accuracy in the one season it "won," the win is telling you about variance, not skill.

The diagnosis: an engine that loves home teams

So why does a reasonable-looking factor model lose? When we audited the picks, the pattern was blunt: roughly 83% of the engine's spread picks were on the home side.

That would be fine if home teams beat the closing spread. They don't. In our data, home teams cover about 49.5% of the time. The market already prices home-field advantage into the line — has for decades — and if anything the modern number slightly over-prices it. An engine that systematically leans home is therefore paying -110 to take the side of a coin that lands its way less than half the time. Stack up 5,490 picks with that lean and a -6.7% ROI isn't bad luck. It's arithmetic.

How did the home bias get in? Death by a thousand plausible constants. A home-field bonus here, a travel-distance penalty on the road team there, an altitude adjustment, a "hostile crowd" factor — each one defensible in isolation, none of them validated against the market, and every single one pushing the adjusted spread in the same direction. The market had already priced all of it. We were adding it twice.

What we changed

The fix was not "tune the constants until the backtest looks better." That's how you overfit your way into a model that wins 2015-2022 on paper and loses 2026 in reality. The fix was philosophical:

  1. Every unvalidated constant went to zero. If a factor never survived our hypothesis pipeline — Welch's t-test, effect size, Bonferroni correction for multiple comparisons, and a temporal split (train on 2015-2022, test on 2023-2025) — it contributes exactly nothing to the adjusted spread. No grandfathered gut feel.
  1. Validated coefficients now live in a JSON file, not in code. Each surviving factor carries the coefficient our published hypothesis work actually measured, with the backtest that justified it. When a factor gets validated, the JSON changes and there's a paper trail. When one fails re-validation, it goes back to zero.
  1. The engine is now allowed to say nothing. A side effect of zeroing the unvalidated constants: the totals model no longer produces a single play, because no validated combination of factors moves a projected total far enough past the 2.5-point edge threshold to justify one. Zero totals picks across 10,886 eligible games. That's not a bug — that's the engine honestly reporting that, on totals, we haven't yet found an edge worth acting on.

We also documented every leakage risk in the backtest itself: rankings are pregame AP only, conference assignments are realignment-aware by season (no 2024 conferences leaking into 2016 games), and no current-season team stats were allowed anywhere near a historical prediction. A backtest that cheats is worse than no backtest, because it manufactures confidence.

Why publish this at all?

Because the alternative is what the rest of the industry does: show you the 2025 season, hide 2015-2024, and call it a track record.

A picks site that only publishes wins is unfalsifiable, and unfalsifiable is another word for useless. The entire value of an analytics operation is that its claims can be checked — that when we say a factor is worth 1.4 points, there's a test behind it, a sample size, a confidence interval, and a correction for how many things we tried before this one worked. Publishing a losing backtest is the cost of admission to that standard. It's also, frankly, the most useful piece of research we produced this offseason: it told us precisely which parts of our own thinking the market had already priced.

The market is very good. Closing lines at major books encode an enormous amount of information, and the honest starting assumption for any model is that it will lose to them at -110. Our backtest confirmed that assumption for our own engine. That's not a failure of the research — that is the research.

The bottom line

Rick's Picks publishes statistical analysis of college football for informational and entertainment purposes. Nothing here is betting advice. 21+. If gambling is affecting you or someone you know, call 1-800-GAMBLER.