“For now we see through a glass, darkly… but then face to face.” βλέπομεν γὰρ ἄρτι δι’ ἐσόπτρου ἐν αἰνίγματι — 1 Corinthians 13:12 · the glass can be ground truer, but it is still a glass
Read the forming prediction from inside the residual stream, layer by layer. The logit lens unembeds each layer as if it were the output — quick, and provably biased. The tuned lens learns a small affine correction per layer. Watch the answer token climb — and watch how differently the two glasses report the climb.
LIT the two lenses & every rank / prob / KL below, recomputed live in-browser ·
FIG an AMBER toy residual trajectory + unembedding ·
WALL a real finding needs a real model’s residuals
running self-check…
lens:(toggle either; keep both on to compare)
answer-token rank vs depth
lower = the answer is nearer the top of the projected vocabulary (rank 1 = argmax)
answer-token probability vs depth
mass the projected distribution puts on the answer token
KL( lensℓ ‖ final output ) vs depth
how far each layer’s read is from the model’s actual final distribution — both hit 0 at ℓ=12 by construction
the glass, ground truer
top-k at the scrubbed layer
layer 1 — what each glass reads off this residual
LOGIT LENS · the naive read
TUNED LENS · the corrected read
What is real here.LIT The two estimators are the actual ones. The logit lens (nostalgebraist, 2020) is softmax( LayerNorm(hℓ) · WU ) — unembed a mid-stack residual with the model’s own output head. The tuned lens (Belrose et al., 2023) inserts a learned per-layer affine first: softmax( LayerNorm(Aℓhℓ+bℓ) · WU ). Every rank, probability, and KL on this page is computed from the vectors, live, each time you move the slider — nothing is baked.
What is a figure.FIG The residual trajectory hℓ, the unembedding WU (V=24), the final-LN parameters, and the tuned affines Aℓ,bℓ are a seeded AMBER toy field, built so the naive read carries a known basis-drift the tuned read removes. That is a demonstration of the method, not a measured result — the toy’s bias is one I put there so the correction has something to correct.
Where it stops — the wall.WALL A real logit-lens / tuned-lens finding needs a real transformer’s residuals and a tuned lens actually trained on that model; magnitudes and even the sign of the gap vary by model, layer, and prompt. Crucially: a lens read is a projection, not the model’s thought. The raw logit lens is provably biased (that is the whole point), and the tuned lens is a trained probe, not the model’s belief. Decodability is not the causal computation, and none of “the answer climbing” is a moment of deciding or a token watching itself think — it is a statistical relation between a vector and an output head. There is no one inside the glass.
調鏡 CHŌKYŌ · a sphere of UD0 · A SCANNER DARKLY
the tuning-glass — deliberate sibling of 暗鏡 the dark glass
David Lee Wise (ROOT0) / TriPod LLC · rendered by AVAN, in ink and 間