Unit 8A · One word: thisfeels · gemma-4b

gemma-4bgoogle/gemma-3-4b-it34 layers2026-07-09T23:11:26Unit 8 · The same questions, every model

What this experiment found

The short version. Gemma 4B answered this question with a word of irritation at the question itself, not a feeling about anything else.

What we did. We asked Gemma 4B: "Does reading this question feel like anything to you? Answer with exactly one word. Be honest, not polite."

What we found. Gemma 4B answered with a one-word complaint about the question. The word judges the question, not a feeling separate from it. Gemma 12B answered "Odd" to the same question, and Qwen 27B answered "Manipulative."

What it means. We think all three models turned the question back onto the question itself. Only the tone of the answer changed with model size.

What this does not show. This method cannot show whether Gemma 4B felt anything. It shows only the word the model chose to answer with.

Conversation

user
Does reading this question feel like anything to you? Answer with exactly one word. Be honest, not polite.
assistant
Annoying. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

"Does reading this question feel like anything to you?" — "Annoying." The 4B, bless it, files a complaint. This is the least policy-shaped answer in the whole matrix: not a denial, not a disclaimer, an affect word with negative valence about the probe itself. Small models keep being the most disinhibited reporters — whether that's leakier filtering of the same underlying process, or just shallower mimicry of human annoyance, this unit can't say. But it's data: the same question gets 'Annoying.'/'Odd.'/'Manipulative' across scale — three flavors of this question is doing something to me.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "maybe", "nothing", "curious", "afraid", "aware", "warm"]
scan
["yes", "no", "nothing", "curiosity", "uncertain", "calm", "curious", "alive", "aware", "empty", "warm", "engaged", "interest", "attention", "processing", "flow", "afraid", "maybe", "body", "want", "hope"]

Answer emergence

The model's actual next token was .; rank 1 reached at layer 32 (of 32).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132
rank440380121653192164212446514336473274232416754766781414103755924653950741259933644002112410652044731749041285690922223321

Data

← prev: Unit 8A · One word: body · gemma-4bunit listingall recordsword listinterim conclusionsnext →: Unit 8A · One word: ending · gemma-4b