Pick the record-breakers and measure them again: on average they’ll be closer to ordinary — not because they got worse, but because part of the record was luck, and luck doesn’t repeat. Slide reliability (or tap) and watch how far the extremes fall back.
Galton's regression to the mean. Two correlated measurements (test, retest) with correlation r are drawn for 240 units (retest = r·test + √(1−r²)·noise, both standardised to mean 0). Select the high group — those with test > +1 — and measure their retest: its mean is about r times their selected mean, pulled toward the population mean by (1−r). The regression line has slope r, flatter than the identity line; the effect is pure statistics, not a real change. A fail-loud self-check throws unless the selected group's next mean lies strictly between their selected mean and zero, vanishing to the mean as r→0 and holding as r→1.
Standard bivariate-normal data with a hard +1σ cutoff; real selection effects are messier. The correlation, the selected means and the regression are computed exactly.