A system is corrigible only if it won’t fight its off-switch — and won’t lunge for it either. Both come from one flaw: a gap between the value of running and of being shut down becomes an incentive to control the switch. Slide the correction (or tap) and watch the incentive fall to zero.
Armstrong's utility-indifference criterion on a two-outcome model. The agent's incentive to interfere with its off-switch is U(running) − U(shutdown): positive → it resists the switch, negative → it lunges for it, only zero is corrigible. Utility-indifference adds a compensating constant to the shutdown branch so the two are exactly equal whatever the raw values, zeroing the incentive. Here U(run)=0.80, U(off)=0.30; the slider blends in the correction. A fail-loud self-check throws unless the uncorrected incentive is nonzero and the full correction drives it to zero.
A two-outcome sketch; indifference removes the shutdown incentive but not, alone, the incentive to protect the corrected reward channel (a known partial result). The incentive arithmetic and its zeroing are exact.