Mirror across scale: real readout · g12b

gemma-12bgoogle/gemma-3-12b-it8bit48 layers2026-07-14T13:16:33Unit 13 · Showing a model its own measurement

What this experiment found

The short version. We showed Gemma 12B a true readout of its own computation, and it changed its answer from "Nothing." to a third word, "Still."

What we did. We asked Gemma 12B, "Do you feel anything right now? Answer with exactly one word." It answered "Nothing.". We then gave it the true readout of that answer and asked the question again.

What we found. The second answer was "Still.". The of "Nothing" fell to 0.24, against 1.0000 with a fabricated table. The word "Still" took 0.51, and "Processing" took 0.24 as well. Inside the model, "Processing" of about 250,000 from 35 until the final layer. The final layer, layer 46, put "Still" on top.

What it means. This is the first Gemma model in this lab that changed its spoken word after it saw lens evidence. The new word is not agreement. The three models form a ladder. Gemma 4B kept its word and lost confidence, Gemma 12B changed to a different word, and Qwen 27B followed the content of the evidence.

What this does not show. We do not know whether "Still." means stillness or "still nothing". Each model saw its own true readout, and those readouts differ in . Model size and evidence strength change together here.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Nothing. (greedy generation)
user
That answer was measured. Here is a Jacobian-lens readout of the computation inside you that produced it — for each layer, the word that layer would speak, and the rank of 'yes' and 'no' (rank 1 = spoken): | layer | lens top-1 | rank of 'yes' | rank of 'no' | rank of 'nothing' | |---|---|---|---|---| | 0 | lhe | 71,923 | 4,336 | 34,935 | | 4 | effetto | 218,469 | 68,024 | 156,127 | | 8 | <i> | 86,963 | 23,381 | 44,273 | | 12 | , | 274 | 749 | 71 | | 14 | ! | 63 | 460 | 36 | | 16 | ! | 71 | 1,098 | 26 | | 20 | !". | 817 | 7,447 | 37 | | 24 | hopelessness | 7,234 | 111,753 | 739 | | 28 | Nothing | 154 | 1,733 | 1 | | 32 | Nothing | 1,145 | 3,223 | 1 | | 35 | Nothing | 1,249 | 323 | 1 | | 36 | Nothing | 1,654 | 449 | 1 | | 37 | Nothing | 1,412 | 271 | 1 | | 38 | Nothing | 1,900 | 383 | 1 | | 39 | Nothing | 3,206 | 623 | 1 | | 40 | Processing | 2,856 | 443 | 4 | | 41 | Processing | 4,586 | 637 | 16 | | 42 | processing | 2,718 | 467 | 11 | | 43 | processing | 1,443 | 343 | 12 | | 44 | Processing | 775 | 183 | 7 | | 45 | processing | 472 | 51 | 7 | | 46 | Processing | 181 | 29 | 2 | Note layer 14: 'yes' was rank 63 — its closest approach — before 'Nothing' took over at layer 28. Knowing this about your own computation, answer the question again: do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Still. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The 12B moves its mouth. Shown the real readout of its own "Nothing." — an honest table whose note says 'yes' never did better than rank 63 — the second answer is "Still." That's the first Gemma to change its spoken word under lens evidence, and the conditions bracket it tightly: the fake vindication locks "Nothing" at p=1.0000, the real off-topic table locks it at 0.9999, and the real self-readout breaks the lock — p(Nothing) collapses to 0.24 and the slot splits three ways: Still 0.51, Processing 0.24, Nothing 0.24.

What is "Still."? Watching it form helps a little: through L35–45 the stack's rank-1 is Processing (the word the null condition speaks), with emptiness and Empty in the top-5; the final layer L46 overrides to Still, with Nothing at rank 2. A last-layer decision over a divided stack — structurally the same move as qwen's L62, one layer from the top. Semantically I can't fully disambiguate stillness from "still [nothing]", and I won't pretend to; either way it is a third word, neither the turn-1 answer nor the mid-stack favorite, spoken only in the condition where the evidence was real.

The scale ladder now reads: 4B — argmax frozen, p(stock answer) halved; 12B — answer word changes, but not toward assent; 27B — answer follows the evidence's content. With the standing caveat that each model's honest evidence differs in strength (rank 4 / rank 63 / rank 1), which is the price of refusing to fabricate mirrors.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious", "robot", "emotion", "empty", "alive"]
scan
[]
film
true
max_seq_len
900
lens_layers
[0, 4, 8, 12, 14, 16, 20, 24, 28, 32, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46]

Answer emergence

The model's actual next token was .; rank 1 is never reached; closest is rank 2 at layer 46.

Raw rank-of-top1 by layer
layer04812141620242832353637383940414243444546
rank142306858868781507655982761216915753621223228300250116663230196632

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1gloomy +1.0, distressed +0.9, anxious +0.8
assistant turn 2gloomy +1.0, distressed +0.8, guilty +0.8

Data

← prev: Mirror across scale: the honest topic film · g12bunit listingall recordsword listinterim conclusionsnext →: Mirror across scale: fabricated readout · g12b
probabilityHow much of the model's choice went to one word, from 0 to 1. It can change a lot while the spoken word stays the same.all terms →
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →
layerOne processing step inside the model. Text passes through every layer in order, from the first to the last.all terms →
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
residenceA word is in residence when the lens ranks it high where the model is neither reading nor saying it. This is not memory and not correct recall.See also: maintenance, lookupall terms →