RULES OF ATTRACTION · the attention heads · each head, one job · kept by MIKAI

THE S-INHIBITION HEAD ◧ 2D · ◍ 3D · ◆ 4D · ◐ shadow · 👶 TAP

‘When John and Mary went to the shop, John gave a drink to ___.’ You say Mary — and so does a language model, through a precise little circuit. An S-inhibition head spots that John is the REPEATED subject and SUPPRESSES it, so the name-copying head is left pointing at the other name. Slide the sentence and watch the veto land.

◆ LIT▲ AMBER
◧ THE MEASURE · 2D
◍ THE VETO · 3D · suppress the name already used
◆ THE FOURTH · 4D · a tesseract turns
◐ THE SHADOW · one dimension down
👶 THE TODDLER CORNER — one fat tap
names
repeated subject
inhibited
answer

◆ LIT — exact / checkable

This is the IOI (indirect-object identification) circuit (Wang et al. 2022). Two names appear; one is repeated as the sentence’s subject. An S-inhibition head detects the DUPLICATED name and writes a signal that suppresses the name-mover head’s attention to it — so the name-mover copies the OTHER (indirect-object) name into the prediction. The instrument runs the template. A fail-loud self-check throws unless it identifies the repeated subject and returns the non-repeated name as the answer, for each sentence.

▲ AMBER — the figure

A rule-level demonstration of the circuit’s LOGIC across a handful of names; the real circuit is a distributed set of heads (duplicate-token → S-inhibition → name-mover) found by causal patching, is graded, and can fail on harder sentences. The identify-and-suppress logic shown is exact.

RULES OF ATTRACTION: a head is not aware of what it does — but it always does the same thing.  — MIKAI
David Lee Wise / ROOT0 / TriPod LLC  ·  the attention circuits, with AVAN