FAITHFULNESS HEAD

Associated with chain-of-thought faithfulness: whether the model’s stated reasoning matches the computation that actually produced the answer. Strengthening it aligns the visible trace with the internal path. Conceptual demo — mechanism real, values illustrative.
AMBER
2D · Quantitative Channel
CANVAS 2D
demo: sine — replace in draw2D()
CANVAS 2D UNAVAILABLE
3D · Spatial Channel
SOFT-GL · NO GPU · NO LIB
DRAG ROTATE · WHEEL ZOOM
demo: lattice + orbit — replace in build3D()/draw3D()
RENDER FAILED —
Notes

Faithfulness head. Linked to whether a model’s chain-of-thought faithfully reflects the computation that produced its answer — the consistency between the stated reasoning and the internal path. Enhancing these heads’ activation has been shown to make CoT traces more consistent with internal behavior. This is the fuzziest of the named types; shown here as a concept (two traces that either coincide or diverge). Amber tier: the mechanism is real in the literature, the alignment values are illustrative.