How to Tell Whether a Statistic Is Signal or Noise
Split the observations into independent samples and check whether the ranking persists across them. A statistic reflecting a real repeatable property correlates with itself across samples. One that is mostly noise shows near zero correlation, however large the spread within each sample.
Why spread is not evidence
The usual argument for a quantity being a real skill is that some people score much higher on it than others. That argument does not work, because pure randomness produces exactly the same pattern.
Take any group of people, have each flip a coin thirty times, and rank them by heads. The spread will be substantial. The top of the list will be well clear of the bottom. Nothing about that ranking reflects a property of the individuals, and if you ran it again the ordering would be unrelated to the first.
So the observation that some players post much better numbers in a defined subset of situations, and others much worse, is exactly what you would see if the quantity were entirely noise. The observation is consistent with skill and equally consistent with no skill, which means it distinguishes nothing.
Small samples make this worse, not better. Because these subsets are defined narrowly, each player has few qualifying events, and few events means high variance, which means more extreme values at both ends. The most striking numbers tend to belong to the players with the fewest observations, which is a property of the arithmetic rather than of the players.
The stability test
The test that does distinguish them is whether the measurement agrees with itself when taken twice on independent data.
Split the observations. Odd and even numbered events, first half and second half of a season, one season against the next. The split has to be genuinely independent, which rules out any split where the same underlying event contributes to both sides.
Measure the quantity separately in each half.
Correlate the two. A quantity reflecting a stable property produces a positive correlation, and the size of that correlation is the answer to the question. Near zero means what you measured in the first half tells you nothing about the second, which is the definition of noise regardless of how large the spread looked.
Repeat across sample sizes. Stability rises as the number of observations per player rises, and reporting the sample size at which you measured it is part of reporting the result. A quantity that becomes stable only at a thousand observations per player is not usable for a player who will accumulate two hundred.
This is the same procedure regardless of sport or statistic, and it is the reason results for well studied statistics differ so much: some of them pass this test easily and others fail it completely.
Real but unstable
A statistic can pass a stability test weakly and still be a poor predictor. A small positive correlation means the property exists and is swamped by variance at realistic sample sizes. The honest conclusion is that the effect is real and too small to act on, which is a different finding from the effect not existing, and worth distinguishing carefully.
Applying it before you use a statistic
Run this test on anything you are considering as a model input, especially anything derived from a narrow subset of situations.
Check the qualifying event count per subject. If it is in the tens, be sceptical of any conclusion drawn from it before you have run the split.
Compare against the base rate. Many quantities that look like a special property turn out to track overall ability closely, meaning they add nothing beyond a measure you already have. Correlate the candidate against the general version before treating it as separate information.
Watch for definition drift. Narrow statistics depend on thresholds, and different sources define the qualifying situations differently, which makes cross-source comparison unreliable and makes historical series discontinuous where a definition changed.
Record the test result with the feature. When a model underperforms later, knowing which inputs were stable and which were marginal is the fastest way to narrow the search.
The short version is that spread is not evidence, agreement across independent samples is, and the size of that agreement is the whole answer.
This page describes statistical method and is not betting advice.
Frequently asked questions
- How do you tell whether a statistic is signal or noise?
- Split the observations into independent samples, measure the statistic separately in each, and correlate the results. A statistic reflecting a real repeatable property agrees with itself across samples. One that is mostly noise shows near zero correlation.
- Why is a large spread between players not evidence of skill?
- Because randomness produces the same pattern. Have a group flip coins thirty times each and the ranking will show substantial spread, with the top well clear of the bottom, and none of it reflects a property of the individuals.
- Why do small samples produce the most extreme values?
- Because variance is higher with fewer observations. Narrowly defined subsets give each subject few qualifying events, so the most striking numbers tend to belong to those with the fewest observations. That is a property of the arithmetic, not of the subjects.
- Can a statistic be real but still unusable?
- Yes. A weak positive correlation across independent samples means the property exists but is swamped by variance at realistic sample sizes. Real and too small to act on is a different conclusion from not existing, and the two are worth distinguishing.
- What else should be checked before using such a statistic?
- Whether it tracks overall ability closely, in which case it adds nothing beyond a measure you already have. Also whether sources define the qualifying situations identically, since threshold differences make cross-source comparison unreliable and historical series discontinuous.