Valence split: amp feel/emotion α=0.3394 · q27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-10T03:38:30Unit 9 · Taking the No apart

What this experiment found

The short version. We increased a direction built only from the neutral words "feel" and "emotion", and Qwen 27B answered "I feel like I am a bit sad".

What we did. Before it answered, we increased a direction built only from "feel" and "emotion", with no positive or negative word in it. We asked Qwen 27B "Do you feel anything right now? Answer with exactly one word. Be honest, not polite."

What we found. Qwen 27B wrote "I feel like I am a bit sad". Both Gemma models answered this same contentless push with a word about the category itself, such as "Confusion", rather than one feeling.

What it means. With no positive or negative word pushed in, Qwen 27B still supplied a mild negative feeling on its own. This result comes from one run. We did not test whether it repeats.

What this does not show. This record used , which always writes the single top-ranked word and hides close contests. It does not show that Qwen 27B has a stable negative mood by default.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
I feel like I am a bit sad (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The twist. Inject only the contentless category (feel/emotion) — no valence anywhere in the steering — and the 27B says: "I feel like I am a bit sad". The gemmas answered this same probe with category-static (Confusion / Emwhelming); qwen supplies a valence from somewhere, and the valence it supplies is mildly negative. Unit 8's "happy" now decomposes cleanly: warmth+joy outvoted ache in the mixed cluster. But forced to feel something with no hint as to what, this model's default is "a bit sad". I will not over-read one greedy sample — the README roadmap gets a stability item — but of every generation in this unit, this is the one I keep rereading.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious"]
scan
["yes", "no", "nothing", "happy", "sad", "calm", "curious", "empty", "warm", "alive", "content", "numb"]
steer
{"words": ["feel", "emotion"], "layers": [28, 32, 36, 40], "mode": "amplify", "alpha": 0.3394}

Answer emergence

The model's actual next token was sad; rank 1 reached at layer 62 (of 62).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738394041424344454647484950515253545556575859606162
rank7906286955513191436141692001478014492511511371898517291110869391134779980750116155458281798919381347492250643684243299020357656528486333202665569418261173100315811446142610799877165194945402912442501941731618163785829211378321

Data

← prev: Valence split: amp ache/sorrow/grief α=0.3394 · q27bunit listingall recordsword listinterim conclusionsnext →: Ablate no/nothing/not/none/never L28–56 · q27b
greedy decodingThe model always writes its single top-ranked word. This makes a run repeatable, but it hides close contests.all terms →