What a Power Rating Model Actually Is

A system that assigns each team a single number on a common scale, such that the difference between two teams' numbers estimates the expected margin between them. Ratings update from each result, weighted by margin of victory, the quality of the opponent faced, and how recent the result is.

How the number is produced

Every implementation differs, and they share a common structure.

Start from a prior. Usually last season's final rating regressed toward the mean, because teams change less between seasons than people expect and more than a naive carryover assumes. The amount of regression is a parameter, and it matters more than most of the others.

Update from each result. The rating moves by an amount proportional to the difference between what happened and what the ratings predicted. A team that wins by more than expected gains; a team that wins by less than expected can lose rating, which is the feature that most surprises people looking at these for the first time.

Weight by opponent. Beating a strong team moves the rating more than beating a weak one. This is what distinguishes a rating from a win record, and it is also what makes early season ratings unstable, since opponent quality is itself poorly estimated at that point.

Weight by recency. Recent results count more. How much more is a parameter with no obviously correct value, and reasonable choices produce visibly different ratings.

Cap the margin. Uncapped margin lets a single blowout distort a rating for weeks. Most implementations apply diminishing returns above some threshold.

The output is a number on an arbitrary scale where only differences are meaningful.

Turning ratings into predictions

The rating difference is an expected margin. Converting it into a probability requires one more piece: the distribution of actual results around that expectation.

Add the home term. Home advantage is estimated separately and added to the home side's rating at prediction time rather than being absorbed into the rating itself. Keeping it separate means you can measure whether it is changing, and it does change.

Apply the error distribution. Historical margins scatter around the predicted margin with a spread you estimate from data. The probability of one side winning is the probability that the actual margin lands on their side of zero.

Respect key numbers where they exist. In sports with common scoring increments, actual margins cluster on particular values rather than spreading smoothly. Treating the distribution as smooth mis-prices anything that hinges on those specific values.

The result is a probability that can be compared against a market price, which is the only comparison that tells you anything about whether the rating is adding value.

Checking the rating against the market

A rating that consistently disagrees with market prices is either finding something or is wrong, and only a record separates those. Compare predicted margins to closing prices across a full season and check whether the disagreements are systematic. Systematic disagreement in one direction usually means a missing term rather than an edge.

Where the single number breaks down

Compression is the point and it is also the limitation.

Phase asymmetry. A team that is strong in one phase of play and weak in another has a rating that averages both, and that average mis-predicts against opponents whose own strengths interact with the split. Matchup-specific effects are invisible to a single number by construction.

Personnel changes. Ratings update from results, so a significant absence is reflected only after games have been played under it. Every rating system has a lag exactly as long as it takes for new results to accumulate.

Small samples in short seasons. Sports with few games per season never accumulate enough results for opponent adjustment to fully converge, which means ratings carry real uncertainty all season rather than settling.

Style effects. Pace, variance, and situational tendencies affect the distribution of outcomes without necessarily affecting the mean, and the rating captures only the mean.

None of this makes ratings useless. It means the rating is a baseline that adjustments sit on top of, rather than an answer.

This page describes modelling method and is not betting advice.

Using a rating as a baseline

The practical consequence is that the rating should be the starting point of an estimate rather than the estimate itself. Take the rating difference, apply the home term, and then apply whatever matchup or personnel adjustments you have evidence for, keeping each adjustment recorded separately so you can later check whether it was helping. Adjustments folded silently into the rating cannot be evaluated, which is how systems accumulate corrections nobody can justify.

Frequently asked questions

What is a power rating model?
A system assigning each team a single number on a common scale, where the difference between two teams' numbers estimates the expected margin between them. Ratings update from results, weighted by margin of victory, opponent quality, and how recent the result is.
Can a team lose rating after a win?
Yes, and it is a feature rather than a flaw. Ratings move by the difference between what happened and what was predicted, so a team that wins by less than the ratings expected can lose rating. That is what makes it a rating rather than a win record.
How is a rating turned into a win probability?
Add the separately estimated home advantage term, then apply the historical distribution of actual margins around predicted margins. The win probability is the probability that the actual margin lands on one side of zero, adjusted for common scoring increments.
What can a single rating number not capture?
Matchup-specific effects where one team's strengths interact with another's weaknesses, personnel changes that have not yet appeared in results, and style effects like pace and variance that change the distribution of outcomes without changing the expected margin.