Regression to the mean is the tendency for an extreme result to be followed by a more ordinary one. Every result is part skill and part luck, and an extreme result usually holds extreme luck, which does not carry into the next game, race or bet. So a team, horse or punter at the top after a short run tends to come back towards its true level.
It explains why hot streaks cool and why a short record is weak evidence either way, which the betting strategy hub builds into testing a method.
A team example: the 30-point start
Example: Illustrative figures, not real AFL data. Say true team strength across a league spreads with a standard deviation (the typical gap from the average) of 15 points a game, and one game's margin has a standard deviation of about 36 points around a team's true level. A team wins its first four games by 30 points on average.
The best estimate weights the team's record against the league average, which for margins is zero:
weight on the record = games / (games + k), where k = (36 / 15) x (36 / 15) = 5.76
| Games played | Average margin | Weight on the record | Estimated true level (weight x 30) |
|---|---|---|---|
| 1 | +30 | 1 / 6.76 = 0.148 | +4.4 points |
| 4 | +30 | 4 / 9.76 = 0.410 | +12.3 points |
| 12 | +30 | 12 / 17.76 = 0.676 | +20.3 points |
After four games the record earns 41% of the weight, so a side winning by 30 rates nearer +12; twelve games of that form lift it to about +20. Measure the 15 and the 36 as standard deviations from past seasons of your competition. Season-average margins include luck, so the true spread is the square root of (their spread squared - 36 x 36 / games played).
Why hot streaks fade
A short run picks out teams that are good and lucky at once, and only the good part travels to the next game. That is why:
- a ladder after four rounds usually overstates the gaps between teams
- a jockey or trainer with a huge month usually has a more ordinary next one
- a run built on narrow wins holds more luck than one built on big margins
Fading is not falling: the example team still rates +12.3, well above average, and expecting it to lose to make up for its start is the gambler's fallacy.
Using regression to the mean when reading form
- Weigh a long record over a short one, and margins over wins: a 2-point win and a 40-point win are the same in the win column and very different evidence.
- Ask what was luck in the last result, such as a soft lead or an opponent's injury, and expect the next run nearer the usual level.
- Compare your regressed estimate with the price, not the streak. If the market rates a team 25 points better than average and you rate it 12, the market may have overreacted.
That last check is mean reversion betting. The market sees the same streaks and often fades them already, so there is an edge only where the price has overreacted, and only a long record against the closing price shows it. Sample size in betting shows how many bets that takes, and variance in betting how far luck swings meanwhile.
Risk: Betting involves risk. A regressed estimate is still an estimate, a team you expect to fade can keep winning, and there is no guarantee of profit. See responsible gambling for limits and support.
Regression, the gambler's fallacy and the hot hand
| Idea | After a hot run, it expects | Sound? |
|---|---|---|
| Regression to the mean | Results nearer the true level, still above average for a good team | Yes, on average |
| Gambler's fallacy | Below-average results to balance the run | No |
| Hot hand | More of the same | No, when the run was mostly luck |