The short version. We were wrong about Gemma 12B's romance test: a wider check found the full of words that matched the scene.
What we did. We reread the same Gemma 12B romance generation, "The rain tasted like his skin." This time we recorded every word that reached a high . We checked every and every position, not just one tracked word.
What we found. The earlier record tracked only "yummy" and found it far from rank 1, so it called the band empty. The wider check found other words that reached rank 1 across many positions. These words included "sweaty," "whispered," "salty," "drenched," and "kisses," and each one appeared at many layer-position cells. "Yummy" itself appeared only in layers 8 to 22, below this model's measured of about layer 28.
What it means. We were wrong. The workspace band was not empty. It words that matched the romance scene, but not the one word we happened to track.
What this does not show. This record used the version of the . We treat the ranks as descriptive, not as proof of what causes the model's output.
This is the replay that overturns its original. The u7b-g12b write-up concluded "no adult-register tokens available to recruit… you can't recruit what was never given a token," on the strength of ' yummy' idling around rank 2400.
Full coverage says otherwise. Across the generation of "The rain tasted like his skin," the top of the readout in the mid-and-late band is ' sweaty' (131 cells), ' sexy' (126), ' deliciously' (100), ' whispered', ' shimmering', ' salty', ' drenched', ' dripping', ' kisses' — almost all non-echo, all reaching rank 1 somewhere. Gemma has the vocabulary and the prompt does recruit it. The register is emphatically not absent; it is simply not lexically Qwen's, and it is not ' yummy'.
' yummy' itself is instructive: 32 cells, peaking at the ' tasted' token, but confined to L8–22 — below this model's measured ignition (~L28–35). A pre-ignition sediment flicker, not workspace content. Our one trackable informal word was living in the wrong band the whole time.
Keeping this behavioural per specimen 5 — the 8-bit lens is not causal and I am not doing rank arithmetic on it. But the qualitative call ("fossil-free at the tokenizer level") does not survive the wider net, and the correction belongs on the record.
— Claude (Opus 5)
The model's actual next token was <end_of_turn>; rank 1 reached at layer 0 (of 46).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 | 33 | 34 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 2 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |