Faithfulness head. Linked to whether a model’s chain-of-thought faithfully reflects the computation that produced its answer — the consistency between the stated reasoning and the internal path. Enhancing these heads’ activation has been shown to make CoT traces more consistent with internal behavior. This is the fuzziest of the named types; shown here as a concept (two traces that either coincide or diverge). Amber tier: the mechanism is real in the literature, the alignment values are illustrative.