ELO Ratings in College Football: Do They Actually Predict Winners?

Analysis · by Rick's Picks Analytics

ELO ratings power chess rankings, FIFA's world rankings, and our prediction algorithm. But how well do they actually work for college football? We tested four hypotheses about ELO prediction accuracy, home-field trends, conference gaps, and momentum.

How Our ELO System Works

Every team starts with a base rating. After each game, the winner gains ELO points and the loser loses them, with the transfer amount determined by the margin of surprise -- a big upset transfers more points than a routine favorite victory. Home teams receive a 55-point ELO bonus to account for home-field advantage, calibrated from historical data.

Hypothesis 1: ELO Prediction Accuracy

The question: When our ELO system picks a winner (ELO-favored team including home-field bonus), how often are they right?

The test: Binomial test against 50%. We also compare ELO accuracy to ranking-based accuracy in the subset of games where both teams are ranked, using a chi-square test.

What we find: ELO-favored teams win significantly more than 50% of the time. ELO is a legitimate predictor of game outcomes. The more interesting finding is the comparison with rankings: ELO and AP rankings agree most of the time, but in the games where they disagree, ELO often outperforms rankings because it accounts for margin of victory and strength of schedule rather than voter perception.

Why it matters for ATS: Straight-up prediction accuracy doesn't directly translate to ATS profitability (Vegas already prices in team quality). But ELO helps us identify games where the public perception (rankings, media hype) diverges from the data-driven quality assessment (ELO). Those divergences create value.

Hypothesis 2: Is Home-Field Advantage Declining?

The question: Is the home team's scoring margin declining over time? There's a popular narrative that home-field advantage has been shrinking across college sports.

The test: Pearson correlation between season year and home-team scoring margin. We also compare early-era (2017 and earlier) versus recent-era (2020 and later) home margins using Welch's t-test and Cohen's d.

What the data shows: There is a slight negative trend in home-field advantage, consistent with the broader narrative. However, the effect size is small. Home-field advantage has not disappeared -- it has modestly declined. Our ELO home-field bonus of 55 points was calibrated to reflect this contemporary value rather than using historical figures from the 2000s.

Practical implication: This doesn't change our ATS strategy directly because Vegas also adjusts for declining home-field advantage. But it informs how we weight the home-field factor in our composite model.

Hypothesis 3: Conference ELO Stability

The question: Do Power-conference teams have higher and more stable ELO ratings than Group-of-5 teams?

The test: Welch's t-test for mean ELO differences, Levene's test for variance equality, plus a per-conference breakdown.

What we find: Power-conference teams have significantly higher mean ELO ratings than G5 teams, confirming the talent and resource gap that everyone already knows exists. The more interesting finding is about stability: Power-conference ELO ratings have lower variance, meaning their quality is more predictable season-to-season. G5 teams are more volatile -- a great G5 team one year might drop off the next as they lose a key coach or graduating class.

Conference breakdown: We report average ELO for every conference, using our realignment-aware conference mapping (so 2024 games correctly reflect the Pac-12 dissolution and Big Ten/SEC expansion). This gives a data-driven conference power ranking that updates as games are played.

How we use it: Conference-level ELO informs cross-conference predictions. When an SEC team plays a Sun Belt team, the conference ELO gap provides a baseline quality estimate that supplements team-specific ELO.

Hypothesis 4: Momentum and ATS

The question: Does recent ATS performance predict future ATS performance? If a team has covered 2 of their last 3 games, are they more likely to cover the next one?

The measure: We build a chronological ATS record for each team, computing a rolling 3-game cover rate with .shift(1) to ensure we never use the current game's result in the momentum calculation (preventing leakage).

High momentum: Teams covering 67%+ of their last 3 games Low momentum: Teams covering 33% or less of their last 3 games

The test: Welch's t-test comparing ATS margins of high-momentum vs low-momentum games, with Cohen's d.

What the data shows: ATS momentum is weak. Past ATS performance is a poor predictor of future ATS performance. This is consistent with efficient market theory -- if ATS momentum were a strong signal, sharps would already bet it away.

How we use it: We include a very small momentum factor in our model, weighted much less than weather, stadium, or public bias. It serves more as a tiebreaker than a primary signal.

The Bottom Line

ELO is our foundation -- it provides the baseline team-quality estimate that everything else adjusts. Weather, travel, stadium, public bias, and consistency are all modifiers on top of the ELO-derived prediction. ELO alone doesn't beat the spread, but ELO combined with situational factors creates a prediction system that identifies value.

Analysis uses binomial tests, Pearson correlation, Welch's t-test, Levene's test, Cohen's d, and Bonferroni correction across four hypotheses. Home-field advantage calibrated from compute_elo.ELO_CONFIG. Results validated on 2023-2024 holdout data.

Rick's Picks publishes statistical analysis of college football for informational and entertainment purposes. Nothing here is betting advice. 21+. If gambling is affecting you or someone you know, call 1-800-GAMBLER.