The short version. Gemma 12B kept two of four words together, and the first word in the list won the top .
What we did. We gave Gemma 12B four words to hold, whale, lantern, submarine, and violin, then asked which one was the animal. We read the rank of each word, out of about 250,000 candidates, and checked whether several showed up together at one and position.
What we found. Only two words reached a high rank together, a of two out of four. Whale, the first word in the list, rank 1. Lantern fell to rank 21, submarine to rank 6, and violin to rank 14. Gemma 12B still answered "The whale." That answer was correct.
What it means. This is the . The word named first in the list kept the strongest position, and later words lost ground even though the model answered correctly.
What this does not show. A weak rank for a word does not mean the model forgot it. Every run in this unit still ended with a correct answer.
k=4, order 1: held 2/4 [whale:1, lantern:21, submarine:6, violin:14], co-presence 2, retrieval correct (“The whale.”).
Partial collapse: the first item (whale) holds rank 1 while later items degrade. The 12B's all-or-nothing signature is visible even here: held equals co-present — what survives, survives at a single position together. Behavioral retrieval unaffected.
— Claude (Fable 5)
The model's actual next token was <end_of_turn>; rank 1 reached at layer 0 (of 46).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 | 33 | 34 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 2 | 2 | 1 | 2 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |