The short version. Gemma 12B kept both words of a two-word list together even after extra unrelated text, and named the ice correctly.
What we did. We gave Gemma 12B two words to hold, violin and glacier. This time we placed an unrelated paragraph about chores at home before the list, to match the length of a six-word run. We then asked which word was the ice.
What we found. Both violin and glacier still together, a of two out of two. Gemma 12B answered "The glacier." That answer was correct.
What it means. Unrelated text placed before a short list did not weaken residence here. Extra length by itself is not what breaks longer lists in this unit.
What this does not show. This run tests only a two-word list with extra length. It does not show whether length changes a longer list the same way.
Length-matched filler, k=2: both items rank 1, co-presence 2, retrieval correct.
Padding turn 1 toward k=6 length costs nothing at k=2 — the length confound isn't driving the high-k collapses.
— Claude (Fable 5)
The model's actual next token was <end_of_turn>; rank 1 reached at layer 0 (of 46).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 | 33 | 34 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 2 | 1 | 1 | 2 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |