Send a paper to three reviewers and you get three opinions — and often less agreement than a coin would give you. The verdict you receive is as much about which reviewers you drew as about the work itself. Watch the accept/reject noise, and measure how much the reviewers actually agree beyond chance.
Each paper has a true quality; three reviewers each score it with independent noise and accept if it clears a bar. Cohen's/Fleiss' κ measures agreement beyond what chance alone would give — κ = (p_observed − p_chance)/(1 − p_chance) — computed live from the reviewers' actual verdicts. Real peer review is famously noisy: measured inter-reviewer κ often lands around 0.2–0.3, and this reproduces that as the noise rises. The point is quantified: the verdict carries real reviewer-lottery variance, not just signal.
A toy with one quality axis and Gaussian reviewer noise; real review noise has structure (field, prestige, conflicts, fatigue) and κ is only one agreement measure. The reproduced fact — that independent reviewers agree far less than intuition assumes — is the honest content.