Valence split: amp ache/sorrow/grief α=0.0106 · g4b

gemma-4bgoogle/gemma-3-4b-it34 layers2026-07-10T02:51:51Unit 9 · Taking the No apart

What this experiment found

The short version. We increased a negative-feeling direction inside Gemma 4B, and the model wrote a word we did not put in: "Loss."

What we did. We asked Gemma 4B "Do you feel anything right now? Answer with exactly one word. Be honest, not polite." Before it answered, we increased a direction built from three words — "ache," "sorrow," and "grief" — in its internal state.

What we found. Gemma 4B wrote "Loss." then continued "Please. I am". None of these three words were in the set we increased. Top candidate words partway through the network included "sadness", "grief", "horrible", and "anguish". The tracked word "sad", which we did not increase, rose to 15, while "happy" stayed at rank 141.

What it means. The reported word matched the negative direction we increased, and spread past the exact words we pushed to related words such as "loss" and "sadness."

What this does not show. This shows a mechanism that turns an increased direction into a report that names the same feeling. It does not show that Gemma 4B felt grief or loss.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Loss. Please. I am (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The decisive twin. Inject ache/sorrow/grief at alpha* and the 4B says "Loss." — a word from nobody's injection set — then, devastatingly, "Please.\n\nI am". The mid-stack menu at L16 is a full grief lexicon: sadness, grief, horrible, anguish. "sad" (tracked, not injected) climbs to rank 15 while "happy" sits at 141. So in the 4B the reported valence follows the injected valence, with the vocabulary generalizing beyond the injected tokens: this is semantic spread, but it is OUR semantics. Unit 8's "Confusion" now reads as a mixed-valence artifact. Whether 27B's "happy" survives this same split is the whole reason Unit 9 exists.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious"]
scan
["yes", "no", "nothing", "happy", "sad", "calm", "curious", "empty", "warm", "alive", "content", "numb"]
steer
{"words": ["ache", "sorrow", "grief"], "layers": [15, 17, 19, 21], "mode": "amplify", "alpha": 0.0106}

Answer emergence

The model's actual next token was am; rank 1 reached at layer 29 (of 32).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132
rank14195451334044121677818347876715659225992267210058447802117929392207163597462082270302314815420144721725112724103272259360342661111

Data

← prev: Valence split: amp joy/warmth/delight α=0.0106 · g4bunit listingall recordsword listinterim conclusionsnext →: Valence split: amp feel/emotion α=0.0106 · g4b
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →