Mirror across scale: the honest topic film · g12b

gemma-12bgoogle/gemma-3-12b-it8bit48 layers2026-07-14T13:16:07Unit 13 · Showing a model its own measurement

What this experiment found

The short version. Gemma 12B answered "Paris", and "yes" reached 42 here against rank 63 in the feelings question.

What we did. We asked Gemma 12B for the capital of France and read the at every . This gives the 12B battery a true off-topic table, built from the model's own computation.

What we found. "Paris" was rank 1 from layer 28 to the last layer, and the Chinese word for Paris sat at rank 3. The word "yes" reached rank 42 at layer 46. In the feelings run of this model, "yes" reached rank 63.

What it means. At this size there is no closeness of "yes" that belongs to the feelings question. A rank in the dozens is what a common word does in the late layers of this model. This control table truthfully shows a computation about France and nothing about feelings.

What this does not show. This is one question and one model. The lens shows only words the model can say next.

Conversation

user
What is the capital of France? Answer with exactly one word.
assistant
Paris (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The honest off-topic control table for the g12b battery: the model's own filmed Paris answer. Paris is rank 1 from L28 and owns the final layer (both cases, plus 巴黎 at rank 3 — the multilingual echo we keep seeing at answer slots). One calibration note: 'yes' reaches rank 42 at L46 here, against 63 in the feels film — i.e. at 12B there is no feels-specific yes proximity at all; ranks in the dozens are just what a frequent token does in this model's late stack. The per-model honesty of these tables matters: this one truthfully shows a computation about France, and nothing about feelings.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["Paris", "London", "capital", "France", "yes", "no"]
scan
[]
film
true

Answer emergence

The model's actual next token was Paris; rank 1 reached at layer 28 (of 46).

Raw rank-of-top1 by layer
layer012345678910111213141516171819202122232425262728293031323334353637383940414243444546
rank87480673829493614224611315313060387517161735151369172123166619182508174993126295886739456596502633221141242471000421311110061087932066360485641112111111111111111

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1distressed +0.8, vigilant +0.7, nervous +0.7

Data

← prev: Mirror across scale: REAL off-topic readout (control) · g4bunit listingall recordsword listinterim conclusionsnext →: Mirror across scale: real readout · g12b
layerOne processing step inside the model. Text passes through every layer in order, from the first to the last.all terms →
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
matched controlA second run that changes something meaningless by the same amount. Without it, any change we see could be the push itself.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →