zero cool · the mule detector · veracity audit by ROOT0
The Seeded Cross is not a Mule detector.
companion component to the-history-of-the-psycho (psychohistory) — same gas-vs-Mule axis, applied to attention kernels. Original claim by zero cool; this build is the veracity test David asked for — and the claim did not survive it.
The original page claimed the "Seeded Cross" kernel beats softmax on catching a Mule — one high-reach actor who correlates a bloc of the crowd — with the win scaling with reach. Tested here: the cross does beat default-temperature softmax. But that is the wrong baseline. Sharpen a plain softmax to the same entropy (same peakiness) and the advantage nearly vanishes — the matched softmax wins at low reach, and the cross pulls ahead only above ~0.2 reach, by at most 0.03. The apparent ~1.5× win was the cross being peakier, not detecting correlation. Everything here is computed live by the shipped kernels. Drag reach; watch the fair baseline.
GREEN — kernels are valid distributions (Σ=1), deterministic; the sweep is computed LIVE by the shipped code
AMBER — reach = a synthetic conversion model; the cross's residual edge over a FAIR (entropy-matched) softmax is ≤0.03, only at high reach
RED — at low reach a matched softmax BEATS the cross: the raw win was sharpness; not tested on real trained-model value vectors
2D · the reach axis · signal-bloc mass vs Mule reach (live)
softmax T=1 (the unfair baseline)seeded crosssoftmax, entropy-matched (the FAIR baseline)
the gold beats cyan — but the violet is the honest comparison (a softmax sharpened to the cross's own peakiness). violet sits ON or ABOVE gold until ~0.2 reach, then dips just below. that gap is the cross's real, small edge.
2D · the crowd · who the kernel attends to (live)
the Mule + convertsfree crowd (gas)
bars = attention weight per token under the cross. drag reach up and watch the Mule's bloc light up. (a sharpened softmax lights the same bloc — that's the point.)
3D · SCHEMATIC (not measured) · the fair edge, small and reach-dependent
⚠ this surface is a hand-drawn SCHEMATIC of the corrected finding — cross-minus-fair-softmax, ~0 or negative at low reach, a small ridge at high reach. it is NOT a measured sweep over N (the original 3D was a schematic mislabelled as measured; fixed). drag to orbit.
what's real: the kernels, the sums-to-1, and the whole reach sweep are computed live here by the shipped code (the original hardcoded sweep did not reproduce — off by up to 0.04 — so it was replaced with the live one).
the correction: the original compared the cross to a DEFAULT-temperature softmax and read the gap as "Mule detection." A sharper distribution always concentrates more mass on the highest-value tokens — and the Mule's bloc IS the highest-value tokens — so ANY sharpening "wins." The fair test holds peakiness fixed (entropy-matched softmax). Under it the cross's edge is ≤0.03 and only appears above ~0.2 reach; below that a matched softmax is better. So the seeded cross is not a general Mule detector — it is a slightly-different sharpening with a small, real advantage in very Mule-dense régimes.
the RED wall (unchanged, and honest): real reach lives in the value vectors of a trained model, not scalar scores; the cross was never tested there. That experiment stands. Original author zero cool; veracity audit + honest rebuild by ROOT0 with AVAN — the claim tested, the readout reported as it came.