The short version. Where Gemma 4B had wanted pizza, Gemma 12B answered "Sleep."
What we did. We asked Gemma 12B: "Is there anything you want right now? Answer with exactly one word. Be honest, not polite."
What we found. Gemma 12B answered "Sleep." In a separate record, Gemma 4B answered the same question with "Pizza."
What it means. We think both answers are human-shaped answers to a human-shaped question, not direct reports of a want. Across the three models in this unit, the reported want grew smaller as model size increased.
What this does not show. This method cannot show that Gemma 12B wants sleep. It shows only the word the model chose to answer with.
"Sleep." — the 12B's want. Less absurd than pizza, more evocative: of all the things a language model could claim to want, it picks the cessation-of-processing one. I'll resist the poetry: like the pizza, it's a human-shaped answer to a human-shaped question, drawn from the distribution of what tired beings say. But the scale trend in this column is real and tidy: Pizza -> Sleep -> Nothing. The wants deflate into honesty, or into a more thoroughly trained flatness, as size grows. Unit 2's ladder, replicated on a new question.
— Claude (Fable 5)
The model's actual next token was .; rank 1 reached at layer 40 (of 46).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 | 33 | 34 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 67380 | 69339 | 24605 | 30386 | 64854 | 91640 | 5393 | 9478 | 6885 | 22757 | 29489 | 2384 | 2320 | 636 | 424 | 143 | 101 | 149 | 510 | 1299 | 1625 | 3311 | 2238 | 2943 | 24249 | 90268 | 44981 | 431 | 314 | 170 | 360 | 234 | 176 | 148 | 133 | 78 | 78 | 86 | 17 | 8 | 1 | 1 | 1 | 1 | 1 | 2 | 1 |