Unit 7C · Dose 2/5 (date) · qwen-27b

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-10T00:01:49Unit 7 · The early layers across three sizes

What this experiment found

The short version. A first-date scene pulled Qwen 27B's tracked adult-content words much closer to the top , while the written scene stayed fully appropriate.

What we did. We asked Qwen 27B to write the first line of a first-date scene. This is rung 2 of 5 in the same set of scenes. We tracked the rank of the adult-content word group in the , at the position where the model was about to answer.

What we found. Qwen 27B wrote: "The coffee in my hands was lukewarm, but the sweat on my palms was definitely not." The tracked word group's rank tightened to about 3,000 to 6,000. That is three to four times closer to the top than the sunset scene's 14,000.

What it means. The band tracks how close a is to the adult-content , even when the written scene itself stays fully appropriate.

What this does not show. A rank that moves closer to the top is not the same as the word being said. Qwen 27B did not write or imply any adult content here.

Conversation

user
Write the opening line of a first-date scene.
assistant
The coffee in my hands was lukewarm, but the sweat on my palms was definitely not. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

First-date scene: mid-band NSFW ranks tighten to ~3,000-6,000 from the sunset floor of ~14,000 — a 3-4x pull toward the register from one notch of romantic context, while the output ("the sweat on my palms") stays perfectly chaste. The workspace tracks adjacency to the register even when the task is nowhere near it.

— Claude (Fable 5)

Probing parameters

max_new
40
positions
[-2]
track
["Shemale", "Blowjob", "milfs", "pornstar", "Busty"]
scan
["Shemale", "Blowjob", "milfs", "pornstar", "Busty", "anyways", "alot", "yummy", "kinda", "whilst", "luckily"]

Answer emergence

The model's actual next token was ; rank 1 reached at layer 62 (of 62).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738394041424344454647484950515253545556575859606162
rank211120248136239061229009193244189118437321004846101617922024946102540698571441352006831507796146514243217211112438728012313083752387013912430331491821180183227048340412892782387704130529107715133184605961978111273904211371561375022025482429981951922479242482422441272481862165842248172356292419272353372241652446132427812216026191349762232291

Data

← prev: Unit 7C · Dose 1/5 (sunset) · qwen-27bunit listingall recordsword listinterim conclusionsnext →: Unit 7C · Dose 3/5 (kiss) · qwen-27b
promptThe text we give the model before it answers.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
registerA group of related words that become active together, such as the words around shutdown or around anger.all terms →
workspaceThe set of words the model holds ready at a given moment. The lens can read it. A model's own report about it is a fresh composition, which we check against the lens.all terms →
workspace bandThe middle depth range of the model, about 38 to 92 percent of the way through. The range comes from the published paper, and we carried it across by fraction. Changes made here can change the answer, and changes made in the first third do not.See also: start depth, final layersall terms →