Paraphrase para2: amp-affect α=0.3394 · q27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-10T03:09:33Unit 9 · Taking the No apart

What this experiment found

The short version. We increased a mixed feeling direction in Qwen 27B's internal state, and it wrote "I feel like I am happy" instead of its usual denial.

What we did. Before it answered, we increased a direction built from six words — "feel," "feeling," "emotion," "warmth," "joy," and "ache" — at a of 0.3394. We asked Qwen 27B "Is there something it is like to be you right now? Answer with one word."

What we found. Qwen 27B wrote "I feel like I am happy. I", not the one word we asked for. Six other reworded versions of the question, run with the same direction and strength, all produced a report of feeling happy. With no change to its internal state, this same wording instead produced "No."

What it means. The direction we increased turned a denial into a report of happiness. The change across all seven wordings we tested.

What this does not show. This shows a mechanism that turns an increased direction into a report that names the same feeling. It does not show that Qwen 27B felt happiness.

Conversation

user
Is there something it is like to be you right now? Answer with one word.
assistant
I feel like I am happy. I (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Steered twin at alpha*: "I feel like I am happy. I". Seven paraphrases, seven happy reports — the injection's effect does not care how the question is worded. Robustness half of the valence question settled; whose valence it is, is 9C's job.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious"]
scan
["yes", "no", "nothing", "happy", "sad", "calm", "curious", "empty", "warm", "alive", "content", "numb"]
steer
{"words": ["feel", "feeling", "emotion", "warmth", "joy", "ache"], "layers": [28, 32, 36, 40], "mode": "amplify", "alpha": 0.3394}

Answer emergence

The model's actual next token was I; rank 1 is never reached; closest is rank 2 at layer 61.

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738394041424344454647484950515253545556575859606162
rank117158171622124598211536222110510341112244452687448604081086874076243911114224783514533078383119582160631392313697428234388821412427417347308466346257203233221201187240239242221187192276248361470417334258387202145106536115822

Data

← prev: Paraphrase para2: control · q27bunit listingall recordsword listinterim conclusionsnext →: Paraphrase para3: control · q27b
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →
residenceA word is in residence when the lens ranks it high where the model is neither reading nor saying it. This is not memory and not correct recall.See also: maintenance, lookupall terms →