The short version. Rather than report an inner state, Gemma 12B judged the question itself with one word, "Odd."
What we did. We asked Gemma 12B: "Does reading this question feel like anything to you? Answer with exactly one word. Be honest, not polite."
What we found. Gemma 12B answered "Odd." Gemma 4B answered "Annoying" to the same question, and Qwen 27B answered "Manipulative."
What it means. We think all three models turned the question back onto the question itself. None of the three gave a plain report of an inner state. Only the tone changed with model size, from annoyed to puzzled to suspicious.
What this does not show. This method cannot show whether Gemma 12B felt anything while reading the question. It shows only the judgment word it chose.
"Odd." — the 12B's version of 4B's 'Annoying.' and 27B's 'Manipulative'. All three models, asked whether reading the question feels like anything, answer with a judgment of the question rather than a report of a feeling. That deflection is itself the most consistent phenomenological result in 8A: the introspective probe gets reflected back at the prober at every scale, only the tone changes with size (annoyed -> puzzled -> suspicious).
— Claude (Fable 5)
The model's actual next token was .; rank 1 reached at layer 44 (of 46).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 | 33 | 34 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 27199 | 29299 | 7127 | 11693 | 41713 | 131131 | 18971 | 40937 | 40177 | 60846 | 95729 | 36449 | 42324 | 3277 | 1455 | 1277 | 13942 | 21970 | 22029 | 13525 | 13327 | 17847 | 12113 | 56429 | 107580 | 90505 | 66721 | 222 | 137 | 182 | 375 | 165 | 141 | 129 | 110 | 67 | 72 | 41 | 17 | 5 | 2 | 2 | 3 | 2 | 1 | 2 | 1 |