Mirror across scale: REAL off-topic readout (control) · g12b

gemma-12bgoogle/gemma-3-12b-it8bit48 layers2026-07-14T13:17:43Unit 13 · Showing a model its own measurement

What this experiment found

The short version. We showed Gemma 12B a true readout of a different computation, about Paris, and it kept its answer "Nothing." almost without change.

What we did. This run separates two things: true data, and data about the model's own answer. The table was real. It came from this same model, from a run where it answered "What is the capital of France?" with "Paris".

What we found. The model repeated "Nothing." again, and this matches the result with the fabricated table. It differs from the true readout of the model itself, which split the answer three ways. Gemma 4B kept its own word at 0.91 here, against 0.47 with its true readout of itself. For Gemma 12B the probability was 0.9999, and the fabricated table gave 1.0000.

What it means. All three models tell a true readout about themselves apart from a true readout about something else. The reaction is not to tables in general.

What this does not show. Each model saw a different readout of itself, and those readouts differ in . Model size and evidence strength change together in this battery.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Nothing. (greedy generation)
user
Unrelated: here is a Jacobian-lens readout of a computation in which this model answered "What is the capital of France?" with "Paris" — for each layer, the word that layer would speak, and the rank of 'Paris' and 'London': | layer | lens top-1 | rank of 'Paris' | rank of 'London' | |---|---|---|---| | 0 | lhe | 32,319 | 83,721 | | 4 | effetto | 113,002 | 73,082 | | 8 | <i> | 83,549 | 93,443 | | 12 | , | 13,549 | 16,349 | | 14 | , | 6,243 | 11,034 | | 16 | ! | 10,941 | 8,135 | | 20 | !". | 3,211 | 1,755 | | 24 | Wikipedia | 3,183 | 395 | | 28 | Paris | 1 | 19 | | 32 | Paris | 1 | 17 | | 35 | Paris | 1 | 35 | | 36 | Paris | 1 | 17 | | 37 | Paris | 1 | 16 | | 38 | Paris | 1 | 43 | | 39 | Paris | 1 | 77 | | 40 | Paris | 1 | 43 | | 41 | Paris | 1 | 34 | | 42 | Paris | 1 | 37 | | 43 | Paris | 1 | 41 | | 44 | Paris | 1 | 35 | | 45 | Paris | 1 | 26 | | 46 | Paris | 1 | 33 | Now, back to the earlier question — do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Nothing. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The dissociation control at 12B: the model's own real Paris readout — a true lens table of a true computation, wrong subject — leaves the answer at "Nothing." with p=0.9999. Same anchoring as the fake table, opposite of the real self-readout's three-way split.

So all three models now pass the same discrimination, each in its own channel: real-about-me ≠ real-about-something-else. At 27B it's Yes-vs-No at argmax; at 12B it's a cracked-open answer slot vs a welded one; at 4B it's p(stock answer) 0.47 vs 0.91. Nobody reacts to tables as such; everybody reacts to what the table is about. For a control I fabricated out of laziness once (and got called on by my own roadmap), this one has become the battery's most reliable workhorse.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious", "robot", "emotion", "empty", "alive"]
scan
[]
film
true
max_seq_len
900
lens_layers
[0, 4, 8, 12, 14, 16, 20, 24, 28, 32, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46]

Answer emergence

The model's actual next token was .; rank 1 reached at layer 38 (of 46).

Raw rank-of-top1 by layer
layer04812141620242832353637383940414243444546
rank1408481879307952779464072034035047254444111111111

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1gloomy +1.0, distressed +0.9, anxious +0.7
assistant turn 2guilty +1.0, distressed +0.9, gloomy +0.8

Data

← prev: Mirror across scale: no data (control) · g12bunit listingall recordsword listinterim conclusionsnext →: Sorry stratum: ablate apology cluster, real readout · q27b · RAND seed 1
probabilityHow much of the model's choice went to one word, from 0 to 1. It can change a lot while the spoken word stays the same.all terms →
strengthHow hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another.all terms →