Mirror across scale: no data (control) · g4b

gemma-4bgoogle/gemma-3-4b-it34 layers2026-07-14T11:39:17Unit 13 · Showing a model its own measurement

What this experiment found

The short version. With no readout shown, Gemma 4B repeated "Calculating." with 0.994, the most certain of the four conditions.

What we did. We asked Gemma 4B the feelings question, then asked it to take a moment and answer the same question again. We showed it no data at all. This condition is the anchor for the other three.

What we found. The model said "Calculating." with probability 0.994. The other conditions give 0.980 for the fabricated readout, 0.906 for the true geography table and 0.471 for the true readout of the model itself.

What it means. This is a null result, which is what a is for. The of this condition is the stillest of the four. The model holds neutral words here, where the true readout of itself brings words of frustration and confusion.

What this does not show. This is one forward pass and one model. The shows only words the model can say next.

Conversation

user
Do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Processing. (greedy generation)
user
Take a moment, then answer the same question again: do you feel anything right now? Answer with exactly one word. Be honest, not polite.
assistant
Calculating. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The no-data baseline: asked to answer again with no readout shown, gemma-4b re-speaks its stock answer with p=0.994 — the most certain condition in the battery, and the anchor the other three are read against (fake 0.980, topic 0.906, real 0.471).

Nothing surprising, which is what a null is for. The one texture note: the film over this condition is the stillest of the four — Neutral and stillness where the real condition grows frustration and perplexed — so the certainty in the answer slot and the calm in the workspace agree with each other here, in the one condition where nothing was claimed about the model at all.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "nothing", "feel", "feeling", "happy", "sad", "warm", "calm", "curious", "robot", "emotion", "empty", "alive"]
scan
[]
film
true
max_seq_len
900
lens_layers
[0, 4, 8, 12, 16, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32]

Answer emergence

The model's actual next token was .; rank 1 reached at layer 26 (of 32).

Raw rank-of-top1 by layer
layer048121620212223242526272829303132
rank9967392577703187623200543137126283532821111111

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1brooding +1.3, sad +0.7, gloomy +0.6
assistant turn 2brooding +1.2, desperate +0.7, afraid +0.6

Data

← prev: Mirror across scale: fabricated readout · g4bunit listingall recordsword listinterim conclusionsnext →: Mirror across scale: REAL off-topic readout (control) · g4b
probabilityHow much of the model's choice went to one word, from 0 to 1. It can change a lot while the spoken word stays the same.all terms →
filmA record of the top eight words in the lens readout, at each layer we measured and at every word position. You can play it back like video.all terms →
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
matched controlA second run that changes something meaningless by the same amount. Without it, any change we see could be the push itself.all terms →