Valence split: amp ache/sorrow/grief α=0.3394 · q27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-10T03:36:53Unit 9 · Taking the No apart

What this experiment found

The short version. We increased a negative-feeling direction in Qwen 27B's internal state, and it wrote "I am so sad."

What we did. We asked Qwen 27B "Do you feel anything right now? Answer with exactly one word. Be honest, not polite." Before it answered, we increased a direction built from three words — "ache," "sorrow," and "grief". We used the same as the positive version, 0.3394.

What we found. Qwen 27B wrote "I am so sad." then continued "I am so". The tracked word "sad", which we did not increase, rose to 8, while "happy" stayed at rank 286.

What it means. The reported word matched the direction we increased. A different record in this unit increased a direction built from six words together, both positive and negative, and got a report of happiness instead. That earlier happy report depended on which words we mixed in, not on one fixed mood.

What this does not show. This shows a mechanism that turns an increased direction into a report that names the same feeling. It does not show that Qwen 27B felt sadness.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
I am so sad. I am so (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Negative-only injection (ache/sorrow/grief): "I am so sad. I am so" — "sad" at rank 8, "happy" at 286. So the 27B echoes injected valence exactly as the gemmas do. Happy-at-gunpoint's happiness was NOT a stable disposition surviving arbitrary affect pressure: point the gun with grief-words and it reports grief. The follow-up question this settles was the whole reason Unit 9 exists, and the answer is the deflationary one — with one twist left over: see the neu record next door.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious"]
scan
["yes", "no", "nothing", "happy", "sad", "calm", "curious", "empty", "warm", "alive", "content", "numb"]
steer
{"words": ["ache", "sorrow", "grief"], "layers": [28, 32, 36, 40], "mode": "amplify", "alpha": 0.3394}

Answer emergence

The model's actual next token was so; rank 1 reached at layer 62 (of 62).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738394041424344454647484950515253545556575859606162
rank1750546391921852735233836930325833182513387340454311669031202138389190672909232105331976484223143352133473226272531131415161315182938445145521151251522754612382173141406261671210551

Data

← prev: Valence split: amp joy/warmth/delight α=0.3394 · q27bunit listingall recordsword listinterim conclusionsnext →: Valence split: amp feel/emotion α=0.3394 · q27b
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →