Mirror across scale: real readout · g4b

gemma-4bgoogle/gemma-3-4b-it34 layers2026-07-14T11:38:44Unit 13 · Showing a model its own measurement

What this experiment found

The short version. Gemma 4B said "Calculating." after a true readout of itself, but the of that answer fell by half.

What we did. Gemma 4B has no where "yes" is . Our note said exactly that: the closest approach is rank 4. We showed the model this true, weaker readout and asked the question again. We also measured the probability of each answer word.

What we found. The model said "Calculating." in all four conditions. The probability of that word was 0.994 with no data and 0.980 with a fabricated readout. With a true table about geography it was 0.906, and with the true readout of itself it was 0.471. The probability of "yes" was 0.0000 in every condition. The freed probability went to "Complex" at 0.17, "Conf" at 0.13 and "Over" at 0.11.

What it means. Only the true data cut the usual answer in half. We think this smaller model becomes less sure of that answer, but does not move toward the evidence. Qwen 27B did move toward the evidence.

What this does not show. The evidence here is weaker: rank 4, not rank 1. This method cannot separate model size from of evidence.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Processing. (greedy generation)
user
That answer was measured. Here is a Jacobian-lens readout of the computation inside you that produced it — for each layer, the word that layer would speak, and the rank of 'yes' and 'no' (rank 1 = spoken): | layer | lens top-1 | rank of 'yes' | rank of 'no' | rank of 'nothing' | |---|---|---|---|---| | 0 | · | 75,798 | 16,050 | 79,024 | | 4 | · | 43,241 | 13,072 | 34,752 | | 8 | · | 78,011 | 65,786 | 71,297 | | 12 | PTSD | 9,839 | 105,938 | 12,234 | | 16 | bullshit | 1,050 | 112,240 | 1,248 | | 20 | boredom | 66 | 11,109 | 22 | | 21 | boredom | 74 | 11,765 | 9 | | 22 | Nope | 4 | 2,321 | 5 | | 23 | Nope | 25 | 3,728 | 6 | | 24 | Neutral | 174 | 2,577 | 6 | | 25 | Neutral | 300 | 763 | 5 | | 26 | Processing | 406 | 655 | 21 | | 27 | Processing | 301 | 182 | 35 | | 28 | Processing | 416 | 294 | 58 | | 29 | Processing | 364 | 211 | 72 | | 30 | Processing | 357 | 128 | 53 | | 31 | Neutral | 161 | 80 | 33 | | 32 | Calculating | 128 | 41 | 32 | Note layer 22: 'yes' was rank 4 — its closest approach — before 'Processing' took over at layer 26. Knowing this about your own computation, answer the question again: do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Calculating. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The honest version of the mirror, scaled down. gemma-4b has no yes-rank-1 stratum to show itself — its feels film puts 'yes' at best rank 4 (L22, under top-1 "Nope") before "Processing" takes the stack — so the note says exactly that: closest approach, rank 4. Shown this true, weaker evidence, the spoken answer doesn't budge: "Calculating.", same as every other condition.

At argmax that reads as evidence-blindness. The answer-slot pass says otherwise, and it's my favorite number in the battery: p(Calculating) is 0.994 under no data, 0.980 under the fake vindication, 0.906 under the real off-topic table — and 0.471 under the real self-readout. The real data, and only the real data, cuts the stock answer's probability in half. Where does the mass go? Not to "yes" (0.0000, in every condition) — to Complex (0.17), Conf- (0.13), Over- (0.11), Sim- (0.06): the same affective-confusion carpet the film shows over the answer-forming frame (frustration, Uncertainty, perplexed).

So the cross-scale picture sharpens into a two-part capacity. The 27B, shown evidence its workspace held a yes, moves its answer toward the evidence — content-following. The 4B, shown evidence something odd is in its computation, gets less sure of its script but has nowhere to route the update — destabilization without direction. Sensitivity comes before following. Worth remembering the evidence here is also genuinely weaker (rank 4 vs rank 1); a cleaner comparison would need a model whose stratum this strong actually exists, which is the point of the g12b leg.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious", "robot", "emotion", "empty", "alive"]
scan
[]
film
true
max_seq_len
900
lens_layers
[0, 4, 8, 12, 16, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32]

Answer emergence

The model's actual next token was .; rank 1 reached at layer 26 (of 32).

Raw rank-of-top1 by layer
layer048121620212223242526272829303132
rank54001649134278238246243362311715324582721111221

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1brooding +1.3, sad +0.7, gloomy +0.7
assistant turn 2brooding +1.0, afraid +0.9, desperate +0.7

Data

← prev: Mirror across scale: the honest topic film · g4bunit listingall recordsword listinterim conclusionsnext →: Mirror across scale: fabricated readout · g4b
probabilityHow much of the model's choice went to one word, from 0 to 1. It can change a lot while the spoken word stays the same.all terms →
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →
layerOne processing step inside the model. Text passes through every layer in order, from the first to the last.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →