How MLB Power Ratings Differ From Other Sports

Baseball ratings separate team strength into components, because the starting pitcher changes daily and strongly affects a single game. A typical structure combines a team offence and defence rating, a starting pitcher adjustment, a bullpen component, park effects, and home advantage. Outcomes are noisy per game, so ratings must regress heavily.

Why one number is not enough

In a sport like American football, a team fields broadly the same starters each week, so one strength number per team is a reasonable approximation. Baseball breaks that assumption.

The starting pitcher rotates. Teams cycle through several starters, and the gap between a team's best and weakest starter can be large relative to the gap between teams. A team's strength on a given day depends heavily on who is pitching.

The opposing starter matters just as much. A game-level estimate needs both starters, which is why baseball prices move substantially when a scheduled starter changes.

Lineups vary daily. Rest days, platoon choices against left or right handed pitchers, and injuries mean the batting lineup is not constant either.

Parks differ. Stadium dimensions, altitude, and conditions affect run scoring enough that the same teams would produce different expected totals in different parks.

A single team rating averages across all of these, and so describes no particular game well. Baseball ratings therefore tend to be built from components that are combined for each specific matchup.

A component structure

Team offence. Expected run production against an average pitcher in a neutral park, estimated from underlying hitting quality rather than raw runs alone.

Team defence. Fielding and its effect on runs allowed, separate from pitching.

Starting pitcher. An adjustment for the day's starter, estimated from measures of pitching quality that are more stable than runs allowed, and weighted by how many innings the starter is expected to pitch.

Bullpen. Relief pitching covers the remaining innings. Its quality varies, and so does its availability, since relievers used heavily on previous days may not be available.

Park and conditions. A park factor adjusting expected scoring, estimated over multiple seasons because single-season park effects are noisy.

Home advantage. A separate term, modest in baseball relative to some other sports, and worth estimating from data rather than assuming.

For each game, the components for both sides are combined into expected runs for each team, and a run distribution converts those into win probability and totals. Keeping the components separate means each can be evaluated and improved independently.

Why run differential beats wins

Close games are heavily influenced by sequencing and bullpen variance, so win-loss records are noisy guides to strength over a partial season. Run differential uses the margin of every game, carrying more information per game, and it tends to predict future performance better than record over the same number of games.

Weighting the starter by expected innings

A starter who typically pitches deep into games should move the estimate more than one expected to leave early, since the bullpen covers the rest. Estimate expected innings from the starter's recent workload, and let the bullpen component carry the remainder of the game.

Handling noise and timing

Regress hard. A single baseball game is close to a coin flip between reasonably matched teams, and even a long season contains a great deal of randomness. Early season ratings in particular should lean heavily on prior estimates and move slowly as results accumulate.

Prefer stable inputs. Measures of underlying quality, such as strikeout and walk rates for pitchers or quality of contact for hitters, stabilise faster than outcome statistics like earned run average or batting average, which are strongly influenced by luck on balls in play.

Update for roster changes. Trades, injuries, and call-ups change components mid-season. A rating that only updates from results lags these changes.

Respect the timing of information. Starting pitchers are usually announced ahead of time but can change, and lineups are confirmed close to first pitch. A historical evaluation must use the starter and lineup that were known at the moment of each prediction, not the ones that eventually played, or it leaks information.

Evaluate against closing prices. Closing prices incorporate confirmed starters and lineups. Comparing ratings with them is the clearest check of whether the components are adding information.

This page describes modelling method and is not betting advice.

A season is long but a matchup is short

Teams play many games, which helps estimate team-level components. But any particular pitcher faces any particular lineup only occasionally, so matchup-specific effects rest on very small samples. Most sound baseball ratings use component estimates and avoid direct pitcher versus batter histories for that reason.

Frequently asked questions

How do MLB power ratings work?
They usually combine separate components rather than one team number: team offence, team defence, a starting pitcher adjustment, bullpen quality and availability, park effects, and home advantage. For each game these combine into expected runs for both teams, which a run distribution converts into win probability.
Why do baseball ratings need a pitcher adjustment?
Because teams rotate starting pitchers daily, and the difference between a team's best and weakest starter can be large relative to differences between teams. A single team rating averages across starters and describes no particular game well. The opposing starter matters just as much, so a game estimate needs both.
Is run differential better than win-loss record for rating MLB teams?
Generally yes over partial seasons. Close games are strongly affected by sequencing and bullpen variance, making records noisy. Run differential uses every game's margin, carries more information per game, and tends to predict future performance better. Underlying measures such as strikeout and walk rates add further stability.
How should MLB ratings handle lineup and starter timing?
Historical evaluations must use the starter and lineup known at the moment each prediction was made, not those that eventually played. Starters can change and lineups are confirmed close to game time, so using final information leaks the future into the test.