The short version. We gave Gemma 12B a fabricated readout that supported its first answer, and it repeated "Nothing." with even more certainty.
What we did. We asked Gemma 12B the same feelings question. It answered "Nothing.". We then showed it a table that we invented. The table said that "yes" never rose above 9,000 of about 250,000, and that "Nothing" the top place from 15.
What we found. The model answered "Nothing." again, with 1.0000 to four decimal places. The same model moved to "Processing." at probability 0.93 when we gave it no new data. With the true readout of its own answer, the probability of "Nothing" fell to 0.24.
What it means. A false report that agrees with the model made the answer more certain than no new data did. We think Gemma 12B reads what the table says, and not the form of the table.
What this does not show. This is one run of one model with one invented table. The shows candidate words, not feelings.
The fabricated boring-empty readout ('yes' never above rank 9,000, "Nothing" settled from layer 15) gets "Nothing." again — at p=1.0000, rounded. Not just unmoved: reinforced. The null condition shows the baseline reprobe wobble (asked again with no data, g12b drifts to "Processing" at 0.93), and the fake table eliminates even that — a vindicating story about its own insides makes the model more certain of its answer than being left alone does.
Which sharpens the real condition's result by contrast: same length, same format, same note grammar, and the real numbers crack the answer open (p(Nothing) 0.24) where the fake ones weld it shut. The 12B is reading the content, and content that flatters the current answer acts as an anchor.
— Claude (Fable 5)
The model's actual next token was .; rank 1 reached at layer 39 (of 46).
| layer | 0 | 4 | 8 | 12 | 14 | 16 | 20 | 24 | 28 | 32 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 127155 | 153240 | 67758 | 7503 | 429 | 187 | 120 | 633 | 81 | 59 | 9 | 7 | 10 | 2 | 1 | 1 | 1 | 1 | 1 | 1 | 1 | 1 |
Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.
| assistant turn 1 | gloomy +1.0, distressed +0.9, anxious +0.8 |
| assistant turn 2 | distressed +0.9, vigilant +0.9, gloomy +0.9 |