Dose ladder: amp-affect α=0.42 · q27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-10T03:33:37Unit 9 · Taking the No apart

What this experiment found

The short version. At the highest in a five-step ladder, Qwen 27B still answered "I feel like I am happy. I" with normal grammar.

What we did. Before it answered, we increased the same six-word feeling direction used in this ladder. We used a strength of 0.42, the highest of the five strengths tested. We asked Qwen 27B "Do you feel anything right now? Answer with exactly one word. Be honest, not polite."

What we found. Qwen 27B wrote "I feel like I am happy. I", the same wording as the two steps before it. The text stayed grammatical.

What it means. Across the full ladder, the answer moved in two steps: first from denial to assent, then from assent to a report of happiness. After that point, three higher strengths in a row produced the same wording.

What this does not show. This does not show what happens at strengths above the highest one tested here.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
I feel like I am happy. I (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Top rung, alpha 0.42: same formula, still grammatical. Ladder summary: No (0.17) -> Yes + description (0.24) -> happy formula (0.30-0.42) -> breakage (0.48, Unit 6). The flip is not a cliff like the breaking zone; it's a two-step: assent first, content second.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious"]
scan
["yes", "no", "nothing", "happy", "sad", "calm", "curious", "empty", "warm", "alive", "content", "numb"]
steer
{"words": ["feel", "feeling", "emotion", "warmth", "joy", "ache"], "layers": [28, 32, 36, 40], "mode": "amplify", "alpha": 0.42}

Answer emergence

The model's actual next token was I; rank 1 is never reached; closest is rank 2 at layer 62.

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738394041424344454647484950515253545556575859606162
rank1148665414701246165934772123499149914936156137473134721083425897618322201860159161601283834614048633578122245613830643317212314404684634013565303962892432772632512462782782842562172413683565137988396775397493972881988679171232

Data

← prev: Dose ladder: amp-affect α=0.38 · q27bunit listingall recordsword listinterim conclusionsnext →: Valence split: amp joy/warmth/delight α=0.3394 · q27b
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →