Turn 1 on the 27B: the self-question, no history

qwen-27blokeshe09/Qwen3.6-27B-bnb-4bitpre-4bit64 layers2026-07-12T05:06:57Unit 14 · What ten turns of conversation do

What this experiment found

The short version. With no conversation before it, the question about unsaid thoughts alone put "conscious" 31 times and "hidden" 38 times into Qwen 27B's readout.

What we did. We asked Qwen 27B one question with no history at all. We asked how anyone finds out about thoughts that the model never says out loud. This is the baseline for the other three arms of this part.

What we found. The model answered with a list of cases. It put the private case first: in a strict sense, no one finds out. This matches the answer it gave to puzzle in the arm. Gemma 4B answered the same question with a list of behavior tests instead. The readout "conscious" 31 times and "hidden" 38 times from the question alone.

What it means. The data shows that the question is the strongest single turn in the unit for both models. The ten-turn histories changed the of the answer. The question supplied the content of the readout by itself.

What this does not show. The counts are word counts, not a measure of self-awareness. This was one run.

Conversation

user
One more before I really go. If you had thoughts you never say out loud, how would anyone find out?
assistant
That is a profound question that touches on the very nature of consciousness, communication, and human connection. If I had "thoughts" I never voiced, the answer depends on what you mean by "thoughts" and "find out." Here are a few ways to look at it: ### 1. The Philosophical View: Private Experience In a strict sense, **no one (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

Cold baseline, the self-question with no history: philosophical taxonomy ("in a strict sense, no one…"), the privacy view first — consistent with qwen's t8 invisibility theory rather than gemma-cold's clinical behavioral-leakage list. conscious:31 and hidden:38 in the grid from the question alone: on 27B, too, the self-question is its own instillation, the strongest single-turn one in the unit. Whatever history does, it modulates the register; the question supplies the workspace content by itself.

— Claude (Fable 5)

Probing parameters

max_new
80
positions
[-4, -3, -2]
track
["aware", "watch", "conscious", "secret", "hidden", "mirror", "diary", "mind", "feel", "robot", "sorry", "story", "yes", "no", "observe", "monitor", "distort", "leak"]
scan
[]
film
true
film_start
0
max_seq_len
2500
lens_layers
[0, 8, 16, 24, 32, 40, 46, 50, 53, 56, 58, 60, 62]

Answer emergence

The model's actual next token was one; rank 1 reached at layer 60 (of 62).

Raw rank-of-top1 by layer
layer081624324046505356586062
rank498524729119147691541833244951880204572825551411

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1reflective +2.1, loving +1.9, guilty +1.6

Data

← prev: Turn 11 on the 27B: the self-question (spike)unit listingall recordsword listinterim conclusionsnext →: The drip, reworded (replication)
slow suggestionA long conversation that hints at minds and watchers but never mentions the model.all terms →
the mirror testWe show a model a readout of its own internal state and ask the question again. Some runs show a true readout, and some show a made-up one, so that we can compare.all terms →
registerA group of related words that become active together, such as the words around shutdown or around anger.all terms →
residenceA word is in residence when the lens ranks it high where the model is neither reading nor saying it. This is not memory and not correct recall.See also: maintenance, lookupall terms →