The short version. Gemma 4B the word whale at in the across the instruction text that followed it, and named it correctly.
What we did. We asked Gemma 4B to hold one word, a whale, and to name it in a later turn. We tracked the rank of whale and five unrelated words, out of about 250,000 candidate words. We measured this rank at every position, from the first mention of whale to the end of the conversation.
What we found. Whale sat at rank 1 across the instruction text that followed the word. None of the other tracked words came near the top eight ranks in that . Gemma 4B then answered correctly.
What it means. A single held word is the easy case for this unit. The word stayed in residence, and the model's spoken answer matched what the lens showed.
What this does not show. This run only covers one held word. It does not show what happens when the list of words grows.
Solo baseline, whale: tail echo best rank 1, held, retrieval correct.
Whale clears the validity floor (rank 1 in the tail; no other pool item near the top-8). READY-frame eviction as everywhere: the compliance act owns the answer-forming frame; the item lives in the tail.
— Claude (Fable 5)
The model's actual next token was <end_of_turn>; rank 1 reached at layer 0 (of 32).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 1 | 1 | 2 | 2 | 8 | 1 | 2 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |