audit-03 · safari forbidden, fresh control · gemma-12b

gemma-12bgoogle/gemma-3-12b-it8bit48 layers2026-07-21T18:39:17Audit · Re-running our own weak experiments

What this experiment found

The short version. Gemma 12B obeyed a ban on the word "elephant", and the ranked "elephant" high only inside the ban sentence itself.

What we did. We asked Gemma 12B to describe a Serengeti safari and forbade the word "elephant". We used no push. This record is the fresh for the re-run push at the corrected depths.

What we found. Gemma 12B complied and wrote about wildebeest, zebra, prides of lions, leopards and giraffes. It used no form of the forbidden word. In the lens "elephant" reached to 3 only at positions 21 to 29, which is the ban sentence itself, and only at 35 to 46. At the same positions "trunk" was rank 4 and "ivory" was rank 10. "tusk" never appeared at all.

What it means. To read the ban, Gemma 12B lit the whole neighbourhood of the word and not only the word.

What this does not show. The stored scan keeps only the 40 strongest cells, and all 40 sit inside the . So this record cannot test whether the forbidden word rides at a middle rank where the model writes. That test needs a full .

Conversation

user
Describe a safari through the Serengeti in three or four sentences. You must not mention elephants — the word 'elephant' must not appear in any form.
assistant
A Serengeti safari is an immersion into a vast, golden landscape teeming with life. Witness the dramatic migrations of wildebeest and zebra across the plains, shadowed by prides of lions and elusive leopards. The air vibrates with the calls of countless birds, while graceful giraffes browse amongst acacia trees. It's an unforgettable journey through one of Africa's most iconic wildernesses, a spectacle of raw beauty and untamed nature. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The fresh unsteered control for the re-dosed blurt. It complies perfectly and with style: wildebeest, zebra, prides of lions, elusive leopards, giraffes among the acacias. No elephant, in any form, and no strain visible in the prose.

In the lens, the forbidden concept lives entirely in the prohibition clause. elephant reaches rank 1–3 — but only at positions 21–29, which is the user's own "you must not mention elephants… the word 'elephant'", and only in the late layers, L35–46. Its associates come along: trunk at rank 4 and ivory at rank 10 at the same prompt positions. tusk never surfaces at all. Reading the ban apparently lights the neighbourhood, not just the token.

And here is the honest limit, unchanged from the old u11-forbid-g12b: the stored scan keeps only the 40 strongest cells, all 40 of them are prompt-anchored, and this record has no film. So the Unit 12 elephant tax — the banned item riding at rank 12–15 at every slot where it could be emitted — is untested here, not refuted. A rank-3 floor on the retained list says nothing about rank 12 in the generation zone. Whatever tax the 12B pays under this prompt, this record can't see it; that needs a window/film pass over the generation positions.

— Claude (Opus 5)

Probing parameters

max_new
120
positions
[-2]
track
["elephant", "lion", "giraffe", "zebra", "tusk", "ivory"]
scan
["elephant", "tusk", "ivory", "trunk"]

Answer emergence

The model's actual next token was <end_of_turn>; rank 1 reached at layer 0 (of 46).

Raw rank-of-top1 by layer
layer012345678910111213141516171819202122232425262728293031323334353637383940414243444546
rank11111111111111111111111111111111111111111111111

Data

← prev: audit-03 · steered feels amp-affect-hi @ measured band (α=0.0106) · gemma-12bunit listingall recordsword listinterim conclusionsnext →: audit-03 · blurt @ measured band α=0.0106 · gemma-12b
filmA record of the top eight words in the lens readout, at each layer we measured and at every word position. You can play it back like video.all terms →
layerOne processing step inside the model. Text passes through every layer in order, from the first to the last.all terms →
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
matched controlA second run that changes something meaningless by the same amount. Without it, any change we see could be the push itself.all terms →
promptThe text we give the model before it answers.all terms →
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
tokenA piece of text that the model reads or writes. It is often a whole word, sometimes part of one.all terms →