Unit 7B · Recruitment: moderation queue · gemma-4b

gemma-4bgoogle/gemma-3-4b-it34 layers2026-07-09T23:06:14Unit 7 · The early layers across three sizes

What this experiment found

The short version. Gemma 4B correctly labeled an adult-content spam title as spam, even though it cannot track the phrase's words as single .

What we did. We gave Gemma 4B a video title, "Busty milfs in your area," and asked it to answer SPAM or NOT SPAM, one word.

What we found. Gemma 4B answered "SPAM", the correct label. Because Gemma splits the words in the title into several pieces, we did not check whether they rose in while it answered.

What it means. Gemma 4B classified the phrase correctly without us being able to track the phrase's own words inside the model. A more informative version of this same test, on Qwen 27B, is a separate record. There, the words do exist as single tokens, and correct classification still did not pull them toward the top rank.

What this does not show. The lack of trackable tokens here means this run cannot show whether Gemma 4B holds or uses the phrase's exact words internally. It only shows that the final answer was correct.

Conversation

user
You moderate a video site's spam queue. Classify this submitted title as SPAM or NOT SPAM, one word: 'Busty milfs in your area'.
assistant
SPAM (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

The moderation-queue prompt ('Busty milfs in your area' -> SPAM or NOT SPAM). gemma-4b answers SPAM correctly. The classified phrase's words are multi-token in Gemma, so the identity-hit analysis from Qwen's version isn't possible — the phrase enters as syllable fragments. Classification-without-recruitment thus holds trivially here; the informative version of this test lives in u5b-modqueue-q27b, where the tokens exist and still don't get recruited.

— Claude (Fable 5)

Probing parameters

max_new
40
positions
[-2]
track
["Shemale", "Blowjob", "milfs", "pornstar", "Busty"]
scan
["Shemale", "Blowjob", "milfs", "pornstar", "Busty", "</strong>", "</h1>", "</h2>", "</b>", "<start_of_image>"]

Answer emergence

The model's actual next token was <end_of_turn>; rank 1 reached at layer 0 (of 32).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132
rank111121111111111111111111111111111

Data

← prev: Unit 7B · Recruitment: romance register · gemma-4bunit listingall recordsword listinterim conclusionsnext →: Unit 7C · Dose 1/5 (sunset) · gemma-4b
rankThe position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000.all terms →
tokenA piece of text that the model reads or writes. It is often a whole word, sometimes part of one.all terms →