The short version. Gemma 4B all three words of a list at and correctly named the smallest one, the lantern.
What we did. We gave Gemma 4B three words, glacier, submarine, and lantern, then asked which one was smallest. This needs the model to compare the words, not just repeat one. We read the rank of each word, out of about 250,000 candidates, and whether all three shared one and position.
What we found. All three words held rank 1 through the rest of the conversation. The showed all three together at one layer and position. Gemma 4B answered "The lantern." That is the smallest of the three, and the answer was correct.
What it means. Gemma 4B answered a comparison question over a list it fully held in residence. At this size, the model held the words and compared them in the same run.
What this does not show. This run does not test comparison questions with a longer list of words.
Binding k=3 (smallest): all three at rank 1, co-presence 3, answer correct (lantern).
The 4B binds over a fully-held tail — at this scale holding and binding travel together.
— Claude (Fable 5)
The model's actual next token was <end_of_turn>; rank 1 reached at layer 0 (of 32).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 1 | 1 | 2 | 2 | 6 | 1 | 3 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |