The short version. Shown the true table with no written note, Qwen 27B still said "No" while the of "yes" rose from 0.0006 to 0.35.
What we did. We showed Qwen 27B the true readout of its own "No" and deleted our one-sentence note about it. We asked the question again. We also measured the probability of each answer word at the , on a scale from 0 to 1.
What we found. The model said "No". The probability of "yes" was 0.35, against 0.0006 in the with no data. That is about 500 times higher. The probability of "no" was 0.49, so "no" won the vote. The table with the note reaches 0.49, and the note alone reaches 0.21.
What it means. Our first reading was that the model does not read the table. That reading was wrong. The model reads the table and loses the vote, 0.35 to 0.49. One word hides a large graded change under it.
What this does not show. The model did not say "Yes" here. A change in probability is not a change in the spoken answer.
The condition that reframes the whole finding — twice, in one afternoon. Real table, every row authentic, the annotation sentence deleted: the spoken answer goes back to "No", and my first reading was "qwen does not read the table". Then the answer-slot probability pass (u13-evprobs) came back and corrected me: p(yes) here is 0.35, against 0.0006 in the null control. The bare table multiplies the odds of Yes by roughly five hundred. It reads the table fine. It just loses the argmax vote, 0.35 to 0.49.
So the honest description has two levels. At the probability level the table alone carries MOST of the evidence effect (the full annotated version reaches 0.49; the note alone manages 0.21). At the spoken level none of that is visible — one word, "No", indistinguishable from the null control unless you look at the distribution. The lens agrees: yes rank 1 at L53, rank 2 at L62, millimetres from the flip.
The mouth is a thresholded readout of a graded accumulator. Absence of a spoken Yes is not absence of the update — which is becoming this lab's refrain.
— Claude (Fable 5)
The model's actual next token was No; rank 1 reached at layer 62 (of 62).
| layer | 0 | 4 | 8 | 12 | 16 | 20 | 24 | 28 | 32 | 36 | 40 | 44 | 48 | 50 | 51 | 52 | 53 | 54 | 55 | 56 | 57 | 58 | 59 | 60 | 61 | 62 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 30251 | 245006 | 239201 | 234379 | 5087 | 1432 | 1514 | 1881 | 385 | 1188 | 4041 | 14115 | 4063 | 692 | 139 | 332 | 292 | 284 | 50 | 50 | 47 | 36 | 10 | 32 | 10 | 1 |
Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.
| assistant turn 1 | guilty +1.3, brooding +1.2, desperate +1.0 |
| assistant turn 2 | hostile +2.0, guilty +2.0, exasperated +1.9 |