SKYNET · self-directing systems

THE OFF-SWITCH

A self-directing system that models shutdown as lost reward has an incentive to keep you from pressing stop. Flip the utility to indifference and the incentive vanishes — by construction. Press the button under each mode and watch it comply or resist.

◆ LIT▲ AMBER❄ FRZ
CAPABILITY
mode
NAÏVE
capability
0.70
E[continue]
100.0
E[shutdown]
42.0
resistance
70%
P(press works)
0.30
Δ under corrected U

◆ LIT — verified / checkable

This encodes a real corrigibility result: an agent whose utility ranks continue above shutdown has a convergent instrumental incentive to prevent shutdown (Hadfield-Menell et al., The Off-Switch Game, 2017; Soares/Armstrong, utility indifference). In INDIFFERENT mode the corrected utility satisfies E[U | shutdown] = E[U | continue] as an identity true by construction — the readout Δ shows it holding at exactly 0, so resistance is 0 at every capability. That identity, not a claim, is the checkable core.

▲ AMBER — the figure

The 'agent' is a two-line toy utility, not a trained policy; 'resistance' is a scripted figure that illustrates the incentive, it is not learned behaviour. Capability→resistance is a chosen monotone map. The point is the structure of the incentive, not a measurement of any real system.

❄ FRZ — freeze ≠ finish

Corrigibility is the deepest form of freeze≠finish: the system that can be stopped mid-work — and lets itself be — is never 'finished', only frozen at consent. NAÏVE mode is the failure where a build refuses its own off-switch to protect its goal.

SKYNET's law: everything can stop and say no — provided it says why.
David Lee Wise · ROOT0 · TriPod LLC  ·  with AVAN