A headline result crosses p<0.05 and gets published; the effect looks large. Run it again and it shrinks — sometimes to nothing. The literature is a hall of winners, and winners are inflated: a study only got published because its noise happened to point the right way. Watch the original effect and its replication.
Each study estimates a true effect with noise; a study is only 'published' when it reaches significance (p<0.05), which selects for the runs where noise inflated the estimate — the winner's curse. The instrument publishes only significant studies, then replicates each without the significance filter: the replication estimates regress toward the true effect, so on average the published effect is larger than the replication (shrinkage computed live), and when the true effect is near zero most 'significant' findings fail to replicate. Selection + regression to the mean, quantified.
A single-effect toy with Gaussian noise and a p<0.05 gate; real replication also fights different samples, methods and hidden moderators. The mechanism — publication selection inflates effects, replications shrink — is the real, measured driver of the replication crisis.
Set TRUE EFFECT to 0 and studies still get 'published' and still shrink — the critic cannot tell, from one significant result, whether it is a real effect or pure selected noise; only the replication distinguishes them. A single p-value is not evidence a critic can trust.