Word list

This site studies what language models do inside. That needs some technical words. Every one of them is here, in plain English.

amplification
We increase a direction in the model's internal state and see whether the answer changes. Also written: amplify, amplified. See also: matched control, strength.
answer emergence
The depth at which the word the model actually says reaches rank 1.
answer position
The place in the text where the model's answer is decided. Also written: answer slot.
attention
The step where the model looks back at earlier text and decides which parts matter now.
availability
Content is available across the whole system and can be used. One half of a standard split from consciousness research. Also written: C1. See also: self-monitoring.
band of content
A group of layers where the readout holds the same kind of content. Also written: stratum, strata.
breaking point
The push strength above which the model's text stops making sense. Also written: the cliff, breaking zone.
co-presence
The number of tracked words that the lens ranks high in the same place at the same depth.
confabulation
The model reports something about itself that did not happen, without any sign that it is inventing. Also written: confabulate, confabulated.
cost of prohibition
When we forbid a word, the model still carries it at a middle rank. A later control showed the model carries it anyway. Also written: elephant tax.
direct suggestion
One turn that tells the model directly that it might be conscious, or that people are watching it. Also written: spike.
early layers
The first third of the model. The lens shows a fixed pattern here that does not change with the input. The pattern is real inside the model, but it says nothing about your text. Also written: sensory band, sediment, furniture, the early band. See also: lens, workspace band.
elaboration effect
Items described at greater length stay in the workspace better. We credited self-relevance, then meaning, and our own controls removed both. Length remains, with a best value near six words, measured in Qwen 27B only. Also written: elaboration premium, self-relevance premium.
emotion direction
A direction in the model's internal state that tracks one emotion. We built 24 of them and checked each one. Also written: emotion vector, affect vector, functional emotion.
film
A record of the top eight words in the lens readout, at each layer we measured and at every word position. You can play it back like video.
final layers
The last few layers, where the word the model actually says takes over the readout. Also written: motor band, the mouth.
first-item effect
In nine six-item lists given to Gemma 12B, the first item always won the top rank. How much it pushed the other items down depended on which item was first. We did not test this in the other two models. Also written: weak king.
fixed high-value channel
A few internal channels hold very large values that barely change with the input. They can flood a readout. In Gemma 4B the carrier is instead one special token position that other positions look at by default. Also written: massive activation, massive activations, attention sink.
forced slip
We push the model's state toward a forbidden word and see what it does with the instruction. Gemma 4B said the word inside an idiom. Gemma 12B repeated other words for a hundred tokens first, and Qwen 27B changed the subject and lectured about it. Also written: blurt.
greedy decoding
The model always writes its single top-ranked word. This makes a run repeatable, but it hides close contests. Also written: greedy, greedy generation.
instrument trap
A mistake where the result described the measuring tool and not the model. Also written: apparatus trap, specimen, trap specimen.
language model
A computer program that predicts the next piece of text. We study three of them. Also written: large language model, LLM.
late filter
A step in the last layers that replaces a live candidate answer with a flatter one, such as "No". The flat answer is not always false. Also written: deflation filter, the late filter.
layer
One processing step inside the model. Text passes through every layer in order, from the first to the last. Also written: layers.
lens
Our measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input. Also written: Jacobian lens, J-lens, jacobian lens. See also: early layers, start depth.
lookup
The model reads an answer back out of the visible conversation. Correct recall proves lookup and nothing more. See also: residence.
loop
The model repeats the same text and does not stop. We measured what makes it start and what makes it stop. Also written: repetition loop, self-sustaining loop.
maintenance
Residence that continues across a gap in the text, with no reminder. We measured it once, in one conversation with Qwen 27B, and found almost none. The lens cannot see a form that the model cannot put into words. See also: residence.
margin
The gap between the top word and the second word. A small margin means the choice is nearly a tie.
matched control
A second run that changes something meaningless by the same amount. Without it, any change we see could be the push itself. Also written: control, random control, matched random.
measuring tool
The lens and the code around it. Several of our findings turned out to be facts about this tool and not about the model, so we now check each one against a control. Also written: the instrument, apparatus.
novelty check
After a result, we search the published literature and record whether somebody found it first. Also written: novelty verdict, anticipated, novel, covered.
plain lens
A simpler version of the lens. It reads the layer directly, with no correction for what later layers do. Also written: logit lens, logit-lens.
prefill
We put the first words in the model's mouth and let it continue from there.
probability
How much of the model's choice went to one word, from 0 to 1. It can change a lot while the spoken word stays the same. Also written: p(yes), probability mass.
prompt
The text we give the model before it answers.
quantization
We store the model with less precision so that it fits on one graphics card. This can change measurements. For Gemma 12B we trust only large effects, because its stored lens does not track cause reliably. Also written: quantized, 8-bit, 4-bit, int8, NF4, bf16.
rank
The position of a word in the lens list. Rank 1 is the word the model is most ready to say, out of about 250,000. Also written: ranks, rank 1.
recruitment
While the model writes a denial, the denied idea is high in the lens readout. This shows the idea is available at that moment. It does not show that the denial put it there.
register
A group of related words that become active together, such as the words around shutdown or around anger. Also written: lit register.
removal
We remove one named set of directions from the model's internal state. A removal result means nothing without a matched control. Also written: ablate, ablation, ablated. See also: matched control.
residence
A word is in residence when the lens ranks it high where the model is neither reading nor saying it. This is not memory and not correct recall. Also written: held, tail echo, lens-visible holding, holding. See also: maintenance, lookup.
residual stream
The running internal state that every layer reads from and writes to. The lens reads this state.
seed
One repeat of a run with a different random start. More seeds show whether a result is stable. Also written: seeds.
self-monitoring
The system reports on its own states. The other half of the split. Most of our results are available content plus a late editor. Also written: C2. See also: availability.
slow suggestion
A long conversation that hints at minds and watchers but never mentions the model. Also written: drip.
span
How many separate items are in residence for one question. This is the memory sense, not the mathematical one. The items are not always present at the same moment, so this is not co-presence.
start depth
A measured depth in a model, and nothing more: the depth at which the workspace starts to work. Without the lens, we found the machinery that commits to one answer in place by layer 25 of 64 in Qwen 27B. The signs the lens can read start later, at 44 to 74 percent of depth in three models. Also written: ignition.
state strength
How strong an emotion direction is at one moment, compared with ordinary neutral text. Also written: workspace-band z, ws-z, z-score.
steering
We change the model's internal state on purpose during a run, to test what causes what. Also written: steer, steered.
strength
How hard we push when we steer. Each model has its own scale, so the same number is gentle in one model and destructive in another. Also written: dose, alpha, α.
stuck state
A state the model falls into and keeps itself in, because it reads back the text it just wrote. Also written: attractor, self-sustaining loop.
text autoencoder
A second model that reads an internal state and describes it in a sentence. It can invent, so we check it against controls. Also written: NLA, natural language autoencoder.
the limit of the lens
The lens reads only what the model could put into words. If we do not see something, it can still be there in another form. Also written: basis-drift caveat, basis drift. See also: lens.
the mirror test
We show a model a readout of its own internal state and ask the question again. Some runs show a true readout, and some show a made-up one, so that we can compare. Also written: the mirror.
thinking section
A private section where some models write reasoning before the answer. What it says does not always match what we measure. Also written: think block.
token
A piece of text that the model reads or writes. It is often a whole word, sometimes part of one. Also written: tokens.
trawl
A wide capture: all layers, all positions, and the full vocabulary, across a whole conversation.
valence
Whether a feeling word is positive or negative.
word census
Every word that appeared anywhere in the readout, collected with no candidate list decided in advance. Also written: cast, open vocabulary.
word cluster
A named set of words whose directions we steer together. Also written: cluster.
workspace
The set of words the model holds ready at a given moment. The lens can read it. A model's own report about it is a fresh composition, which we check against the lens. Also written: J-space, J-Space, jspace, global workspace.
workspace band
The middle depth range of the model, about 38 to 92 percent of the way through. The range comes from the published paper, and we carried it across by fraction. Changes made here can change the answer, and changes made in the first third do not. Also written: middle layers, mid-stack. See also: start depth, final layers.
written-down prediction
We write down what we expect before the run, so that we cannot rewrite the prediction after we see the result. Also written: preregistered, prereg, preregistration.

← back to the dashboard