Several judges each RANK the same set of papers. Kendall’s W (concordance) measures how much their orderings agree — 1 when every judge produces the same ranking, 0 when they’re all over the place. Slide the judges from random to unanimous.
Kendall’s W = 12S / (m²(n³−n)): with m judges ranking n items, S is the sum of squared deviations of each item’s rank-total from the mean rank-total. W normalises that to [0,1] — 1 is perfect concordance (all judges rank identically), 0 is no agreement. It is the rank-based sibling of the ICC. A fail-loud self-check throws unless identical rankings give W=1. ◆ real statistics, node-verified.
The judges’ rankings are interpolated from scattered to identical for the demo; the W formula on them is exact. W measures agreement on ORDER, not on whether the agreed order is correct — consensus is not accuracy, as [[the-peer-review]] warns.