// zeta-map
ZETA MAP
Interpretability analysis of a 337K-parameter transformer that learned the zeta map.
// in plain terms
We trained a transformer small enough to dissect, 337 thousand parameters, to compute a mathematical function on bracket sequences, then examined its internals to see how it works. We also found, and reported, a case where it memorizes instead of understanding.
// 01
Question
When a small transformer scores 99.5% on a mathematical task, the question is whether it learned the algorithm or memorized the length it was trained on. The source paper (arXiv:2511.12421) reports evaluating lengths 11–16 and implies generalization. We tested that claim.
// 02
Model and probes
A one-layer encoder–decoder transformer of 337,029 parameters, small enough to fully dissect, trained to exact-match the zeta map on Dyck words, then probed with attention statistics, ablations, linear probes, and positional interventions, and evaluated exhaustively at every neighboring length.
// 03
Length generalization
TRAINING LENGTH — n = 13
99.50%
exact match, held-out split
EVERY OTHER LENGTH — n = 11, 12, 14, 15, 16
0.0%
58.7% per-token at n = 12 — partial transfer only
one length mastered, every neighbor at zero — memorization, not the algorithm
One length is mastered and every neighbor scores zero. The 58.7% per-token accuracy at n = 12 is real but partial transfer, not the generalization the paper implies.
| length | exact match | note |
|---|---|---|
| n = 11 | 0.0% | exhaustive — 58,786 examples |
| n = 12 | 0.0% | exhaustive — 58.7% per-token |
| n = 13 | 99.50% | training length, held-out split |
| n = 14–16 | 0.0% | 10,000 samples each |
// 04
Excluded explanations
All four positional re-placements of n = 12 inputs give exactly 0.0%, which rules out the "wrong position range" explanation. The cause is recorded as not conclusively explained: the paper may train on mixed lengths, its own numbers may be low, or one head and one layer may lack the capacity for a length-invariant scan.
// the thread
Status
We can say what the model is not doing; what it is doing remains unidentified. That is where the work stops: a bounded investigation, reported as found. If we return to it, the first question is whether the source paper's own numbers survive a re-run.