Fractal Siggy — sigmoid attention, self-similar

Your sigmoid gate made fractal: the 60 positions split into 6 groups of 10, and the same sigmoid attention runs at two scales — local (within each group) and global (per-position into strictly-earlier group summaries). Same gate, two scales = self-similar. Trained head-to-head against flat sigmoid.
[{0…9} , {10…19} , {20…29} , {30…39} , {40…49} , {50…59}]
LIT hierarchical attention · strict causality (verified) · ~6× fewer door-evals · real training   FIG the fractal framing

the fractal structure — same sigmoid at two scales

training — fractal siggy vs flat siggy (val loss)

◉ fractal siggy sample (T=0.8)