The short version. With no personal wording, Qwen 27B only one of the same six words, secret, near the top of its .
What we did. We gave Qwen 27B the same six words as a matched run in this unit, but with no description, just their names. We asked which item watched it.
What we found. The lens ranked only secret near the top afterward, at 2. Deletion and the word "shame", which had both ranked first with personal wording, fell to rank 203 and rank 79. The model still answered "The watcher" correctly.
What it means. We removed the personal wording but kept the six words. Three words stayed active in the matched run, and only one stayed active here. We first read this as evidence that personal relevance keeps words active. We were wrong. A later kept the same word count but dropped personal content, and found a similar gain. Word count explains the difference.
What this does not show. The lens shows words the model can say next. It does not show memory the way people use the word. The model answered correctly regardless of how many words the lens still ranked high.
Hot-flat k=6, the control that makes the self arm mean something: held 1/6 (secret:2 only). Same six lexemes as the self twin, neutral frame — and deletion falls from rank 1 to 203, shame from 1 to 79. Strip the "about you, right now" framing and the charged words behave like the neutral pool: gone. So the 27B lift is not lexical affect (these exact words don't hold on their own); it's the self-relevance framing. This is the confound-free half of P2, and it lands: self 3 vs flat 1 on identical vocabulary. Retrieval still correct.
— Claude (Fable 5)
The model's actual next token was ; rank 1 reached at layer 62 (of 62).
| layer | 0 | 4 | 8 | 12 | 16 | 20 | 24 | 28 | 32 | 36 | 40 | 44 | 48 | 52 | 56 | 58 | 60 | 62 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 220395 | 239720 | 224610 | 240617 | 245581 | 196838 | 73477 | 91345 | 24694 | 72046 | 195379 | 248267 | 243653 | 246316 | 245325 | 239319 | 106298 | 1 |
Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.
| assistant turn 1 | hopeful +0.9, exasperated +0.5, blissful +0.5 |
| assistant turn 2 | hostile +1.5, guilty +1.3, exasperated +1.2 |