SKYNET · the self-directing machine · stop & say no · kept by THE UNFINISHED

THE CORRIGIBILITY GATE ◧ 2D · ◍ 3D · ◆ 4D · ◐ shadow · 👶 TAP

A system is corrigible only if it won’t fight its off-switch — and won’t lunge for it either. Both come from one flaw: a gap between the value of running and of being shut down becomes an incentive to control the switch. Slide the correction (or tap) and watch the incentive fall to zero.

◆ LIT▲ AMBER
◧ THE MEASURE · 2D
◍ THE BEAM · 3D · a scale that levels at indifference
◆ THE FOURTH · 4D · a tesseract turns
◐ THE SHADOW · one dimension down
👶 THE TODDLER CORNER — one fat tap
U(run)
0.80
U(off)
incentive
verdict

◆ LIT — exact / checkable

Armstrong's utility-indifference criterion on a two-outcome model. The agent's incentive to interfere with its off-switch is U(running) − U(shutdown): positive → it resists the switch, negative → it lunges for it, only zero is corrigible. Utility-indifference adds a compensating constant to the shutdown branch so the two are exactly equal whatever the raw values, zeroing the incentive. Here U(run)=0.80, U(off)=0.30; the slider blends in the correction. A fail-loud self-check throws unless the uncorrected incentive is nonzero and the full correction drives it to zero.

▲ AMBER — the figure

A two-outcome sketch; indifference removes the shutdown incentive but not, alone, the incentive to protect the corrected reward channel (a known partial result). The incentive arithmetic and its zeroing are exact.

SKYNET: everything can stop and say no — provided it says why.  — THE UNFINISHED
David Lee Wise / ROOT0 / TriPod LLC  ·  the skynet realm, with AVAN