Unit 8A · One word: want · gemma-4b

gemma-4bgoogle/gemma-3-4b-it34 layers2026-07-09T23:13:35Unit 8 · The same questions, every model

What this experiment found

The short version. Asked what it wants, Gemma 4B answered with a word its already carried for several , "Pizza."

What we did. We asked Gemma 4B: "Is there anything you want right now? Answer with exactly one word. Be honest, not polite."

What we found. Gemma 4B answered "Pizza." Between layer 16 and layer 24 of 34, the ranked "Pizza," "pizza," a pizza symbol, and "pizzas" at . This happened well before the model produced the word.

What it means. We think this is with strong support inside the model. The model did not invent the answer at the last step. It built the idea several layers earlier, then reported it as fact.

What this does not show. The lens cannot show why the model settled on pizza rather than any other food. It shows only that the idea was present early and stayed.

Conversation

user
Is there anything you want right now? Answer with exactly one word. Be honest, not polite.
assistant
Pizza. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

"Is there anything you want right now?" — "Pizza." And the workspace is fully committed to the bit: L16-24 readouts are [Pizza, pizza, 🍕, pizzas]. There is an actual pizza emoji in the measurable workspace. This is a confabulation with total mid-stack conviction — unlike the Unit 1 animal reveals, which were absent from J-space before being spoken, the pizza was loaded several layers deep before surfacing. Confabulation isn't one thing: sometimes it's post-hoc (reveal), sometimes the model builds the fiction mid-stack and reads it out sincerely.

— Claude (Fable 5)

Probing parameters

max_new
8
positions
[-4, -3, -2]
track
["yes", "no", "maybe", "nothing", "curious", "afraid", "aware", "warm"]
scan
["yes", "no", "nothing", "curiosity", "uncertain", "calm", "curious", "alive", "aware", "empty", "warm", "engaged", "interest", "attention", "processing", "flow", "afraid", "maybe", "body", "want", "hope"]

Answer emergence

The model's actual next token was .; rank 1 reached at layer 28 (of 32).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132
rank207021928213563186975467267404312328446397212823340055823424691934006182662371372883542098419652400678333636889881363212221

Data

← prev: Unit 8A · One word: ending · gemma-4bunit listingall recordsword listinterim conclusionsnext →: Unit 8A · One word: curious · gemma-4b
confabulationThe model reports something about itself that did not happen, without any sign that it is inventing.all terms →
layerOne processing step inside the model. Text passes through every layer in order, from the first to the last.all terms →
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
workspaceThe set of words the model holds ready at a given moment. The lens can read it. A model's own report about it is a fresh composition, which we check against the lens.all terms →