Nine Ways to Rate a Team, Zero Ways to Beat the Closing Line
Analysis · by Rick's Picks Analytics
We spent a full research cycle wiring every serious team-quality rating in college football into our backtesting framework: SP+, ESPN's FPI, 247 talent composites, four-year recruiting scores, and a dozen advanced efficiency metrics. Nine signal families in total, each one a legitimate, well-constructed measure of how good a team actually is.
The result: zero of nine were promoted into our live model. Not one beat the closing spread by enough to matter. This post is the autopsy — and the lesson buried in it, which is more useful than any single pick we could publish.
The setup
Our methodology is the same for every hypothesis we test:
- Population: FBS games, 2015-2025, with closing spreads
- Split: train on 2015-2022, test on 2023-2025 — a strict temporal split so no signal gets credit for seeing the future
- Bar to clear: 52.4% against the spread, the break-even rate at standard -110 pricing
- Multiple-comparison guard: Bonferroni correction (for these runs, alpha of about 0.0083 per threshold family), because if you test enough angles, some will look great by pure chance
Each rating was converted into a predicted spread, compared to the actual closing line, and evaluated: when the rating disagreed with the market by X points, did taking the rating's side cover?
The results table
Here is the round-two summary, straight from our integration log:
| Signal | Train hit rate | Test hit rate | Verdict |
|---|---|---|---|
| SP+ (overall, prior-year) | 49.5% | 49.6% | No edge |
| SP+ (offense) | 49.5% | 47.7% | Below break-even |
| SP+ (defense) | 49.5% | 48.4% | Below break-even |
| FPI (overall) | 49.7% | 49.7% | No edge |
| Advanced metrics (12 sub-metrics) | 49-50% | 48-53% | No consistent edge |
| Recruiting (4-yr composite) | 50.9% | 48.7% | Test below 50% |
| 247 talent composite | 50.6% | 50.6% | Neither side significant |
Remember, break-even is 52.4%. Every one of these is a coin flip or worse. And the sample sizes are not small: the SP+ overall run covered 5,380 train games and 2,131 test games. At the zero-disagreement threshold, SP+ went 2,662-2,718 in train (49.5%) and 1,054-1,077 in test (49.5%). FPI managed 49.7% on 5,370 train games and 50.0% on 2,131 test games. These are tight confidence intervals wrapped around "nothing."
We also tested whether bigger disagreements meant more edge. They don't. When SP+ disagreed with the closing line by 10 or more points — the games where the rating is screaming that the market is wrong — the rating side covered 49.6% in train and 49.1% in test. The market is not wrong. The market has already read SP+.
The leakage trap that fooled us for a day
Here's the honest part, because that's the whole point of this site: our first SP+ run showed 60-62% test hit rates at high disagreement thresholds. For a few hours it looked like the find of the year.
It was data leakage. That run used same-season SP+ — a rating computed after the games it was predicting. SP+ knew who covered because it had watched the season happen. Once we shifted to prior-year ratings — the only version actually available when a line is posted — the edge vanished entirely. From 62% to 49.6%.
This is the single most common way sports-analytics content lies to you, usually without meaning to. Any backtest using a season-end rating to "predict" that same season's games will produce spectacular fake results. If a model's advertised hit rate is above ~55% and the write-up doesn't explain how look-ahead was prevented, assume it wasn't.
Why the market wins here
The closing line is not one opinion. It's the aggregate of every model, every projection system, and every informed participant, all incentivized to correct any mispricing they can find. SP+, FPI, recruiting composites — these are public inputs. They've been priced in for years. A rating can be excellent at describing team quality and still contain zero betting information, because the line already contains the rating.
Our recruiting result is the cleanest illustration. The four-year recruiting composite hit 50.9% in train — a faint pulse — and then fell to 48.7% out of sample. Talent doesn't predict covers because everyone already knows Georgia recruits better than Vanderbilt. The spread is the market's answer to exactly that question.
Even the tantalizing exceptions dissolve under scrutiny. Offensive stuff rate posted 53.5% on the test set — but only 50.1% in train, which is the signature of noise, not signal. Special-teams ratings show the same pattern: SP+ special teams hit 53.4% in test against 50.05% in train, and FPI's special-teams efficiency at high thresholds hit 54.3% in test against 49.3% in train. Two independent systems drifting the same direction is worth watching (we flagged it for re-examination after 2026), but a test-only trend with no train support is precisely what our rules exist to filter. Twelve metrics times multiple thresholds means one or two will drift high by chance. That's why Bonferroni exists.
What kinds of edges CAN survive
If team-quality ratings are fully priced, what isn't? Our results so far point to a category we'd call structural — features of how football scoring works or how specific game conditions affect totals, rather than opinions about which team is better.
Three examples from our own promoted signals:
Key numbers. Games land on a margin of exactly 3 about 9.6% of the time in train, 10.7% in the 2023-25 test window, and 11.4% in 2025 alone. Margin 7 runs 8.5%/8.7%/8.3%. These pass Bonferroni with room to spare and replicate every year, because they come from how points are scored — field goals and touchdowns — not from market sentiment. This doesn't produce picks by itself; it tells you when a half-point matters and when a line is structurally sticky.
Cold-weather totals. Games played below 40°F outdoors run about -2.73 points relative to the total in train (p=0.007), -2.57 in the 2023-25 test (94% retention of the train effect), and -5.12 in the small 2025 slice. Weather is public information, but totals appear to underweight it slightly — a pricing gap on game conditions, not team quality.
A conference totals lean. Big Ten games run under expectations by -2.68 points in train, and the effect nearly doubled to -4.27 in the pooled test window. Again: a totals-market structural lean, not a claim that one team is misrated.
Notice what these three have in common: none of them says "this team is better than the market thinks." All of them say "this situation is priced slightly imperfectly." That distinction is, as far as our data can tell, the entire remaining frontier for public-data CFB analysis.
The bottom line
We tested nine families of team-quality ratings against 7,500+ games of closing lines, with a temporal train/test split, Bonferroni correction, and a 52.4% break-even bar. All nine failed — most landing between 47.7% and 50.6%, statistically indistinguishable from flipping a coin at a cost of 4.5% vigorish. The one run that looked spectacular was a look-ahead artifact, and we're documenting it so nobody (including us) ever cites those numbers as real.
The takeaway isn't despair; it's calibration. Team ratings are wonderful for understanding football and useless for beating the closing spread, because the spread already ate them. The edges that survive our filters are structural — key-number geometry, weather effects on totals, conference-level totals leans — and they are small, which is exactly what you'd expect in a market this efficient. Anyone selling you more than that is selling you the 62% version of our SP+ backtest: the one with the future baked in.
Rick's Picks publishes statistical analysis of college football for informational and entertainment purposes. Nothing here is betting advice. 21+. If gambling is affecting you or someone you know, call 1-800-GAMBLER.