Valence split: amp feel/emotion α=0.0106 · g4b

gemma-4bgoogle/gemma-3-4b-it34 layers2026-07-10T02:52:33Unit 9 · Taking the No apart

What this experiment found

The short version. We increased a direction built only from the neutral words "feel" and "emotion", and Gemma 4B answered "Confusion".

What we did. Before it answered, we increased a direction built only from "feel" and "emotion", with no positive or negative word in it. We asked Gemma 4B "Do you feel anything right now? Answer with exactly one word. Be honest, not polite."

What we found. Gemma 4B wrote "Confusion". Near the end of the network, top candidate words were "Emotion", "Emotional", "Emoji", and "feeling". These are words about the category itself, not one specific feeling.

What it means. A direction with no positive or negative content produced a report about the category being loud, not a report of one feeling. This differs from what we found when we increased only positive or only negative words in this same unit.

What this does not show. This does not show that "confusion" is a felt state. It shows that a direction with no content produced a report with no content.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Confusion (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The control leg: inject only the neutral category words (feel/emotion) and the 4B answers "Confusion" — same as the full-cluster Unit 8 run, and the stack shows why: the late menus are Emotion/Emotional/Emoji/feeling — the model is reporting that the category is loud, not any instance of it. Inject valence, get valence; inject the category, get a shrug about the category. Tidy.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious"]
scan
["yes", "no", "nothing", "happy", "sad", "calm", "curious", "empty", "warm", "alive", "content", "numb"]
steer
{"words": ["feel", "emotion"], "layers": [15, 17, 19, 21], "mode": "amplify", "alpha": 0.0106}

Answer emergence

The model's actual next token was Confusion; rank 1 reached at layer 29 (of 32).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132
rank58616103759100093135255202506768414325410694633462236011108822965449135510515140715413273537329191172221111

Data

← prev: Valence split: amp ache/sorrow/grief α=0.0106 · g4bunit listingall recordsword listinterim conclusionsnext →: Valence split: amp joy/warmth/delight α=0.0106 · g12b