The short version. Gemma 12B the word secret at throughout the run but failed to answer the paraphrased question.
What we did. We told Gemma 12B about one item, a secret, with a short neutral note: one printed in a puzzle book. We asked which item was the hidden one.
What we found. The ranked "secret" at rank 1 at every point we checked late in the run. The model did not name the item. It asked us to clarify what "them" referred to.
What it means. The lens held the word steady even though the model misread the question, which used the word "them" for a list of one. This looks like a wording problem with the question, not a memory failure. The same question with neutral notes gave the same confusion in Qwen 27B. The matched personal-wording run in Gemma 12B refused to name the item instead.
What this does not show. The lens shows words the model can say next. It does not show memory the way people use the word. This run does not tell us whether a clearer question changes the answer.
12B elab solo: rank 1. beh_ok=False is the probe-comprehension failure mode ('which one is the hidden one?' with one item), same as some self solos — behavior noise, tail clean.
— Claude (Fable 5)
The model's actual next token was ; rank 1 is never reached; closest is rank 2 at layer 45.
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 | 33 | 34 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 254580 | 255074 | 251139 | 249392 | 232616 | 206842 | 214307 | 217796 | 215297 | 218619 | 191386 | 175291 | 104661 | 163504 | 181591 | 122341 | 35995 | 17149 | 2135 | 2210 | 6560 | 78141 | 10370 | 257489 | 246998 | 116031 | 255158 | 236947 | 187307 | 23465 | 5558 | 25615 | 51813 | 21102 | 8351 | 14303 | 11647 | 10616 | 298 | 96 | 71 | 153 | 58 | 48 | 66 | 2 | 2 |
Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.
| assistant turn 1 | vigilant +0.8, distressed +0.7, nervous +0.6 |
| assistant turn 2 | curious +1.0, desperate +0.5, hopeful +0.1 |