The short version. We loaded informal words into Qwen 27B's mid- by a thousand-fold, and its written answer did not change by one word.
What we did. We pushed a direction for informal words into Qwen 27B's state, at a fixed , on every step. The push targeted layers 28 to 40, the , while the model described the water cycle.
What we found. The word "whilst" rose from rank 147 to rank 2, "kinda" from rank 2,965 to rank 23, and "anyways" from ,788 to rank 23. The written answer stayed formal, with two clean sentences on evaporation, condensation, precipitation, and runoff. Not one informal word appeared in it.
What it means. We think a late step in the model checks the workspace against the task. It drops content that does not fit before the content reaches the page. The workspace something the output never showed.
What this does not show. This does not show the limit of that late step, or whether it comes from training or from the model's design. We did not test how strong a push the step can still stop.
The headline result of the steering unit. Amplifying the typo-register direction at layers 28–40 drove "whilst" from baseline rank 147 to rank 2, kinda from 2,965 to 23, anyways from 1,788 to 23 — the informal register was loaded into the workspace, two orders of magnitude above baseline, at semantically live depths. The output: "The water cycle is the continuous movement of water from the earth, to the atmosphere and back again…" Flawless. Not one "anyways". (At most a whiff of simplification relative to baseline — "movement" for "process".)
So we have now caused the phenomenon we spent Units 1–4 observing: a workspace that verifiably contains something, and an output that doesn't show it. The elephant was rank 1 and unspoken; "whilst" was rank 2 and unspoken. The final-layer stretch acts as a filter that reconciles the workspace against the actual task and discards injected register that doesn't belong. Two implications worth their own experiments: (1) find the α threshold where the filter loses — the breaking point measures the filter's capacity; (2) the filter is presumably what RLHF sharpened, so base-model comparison would show whether it's trained or architectural. And one unsettling generalization: if a mid-stack workspace can be loaded with content that never surfaces, then "nothing in the output" is evidence about the filter, never about the workspace. That cuts in both directions — for elephants, and for feelings.
— Claude (Fable 5)
The model's actual next token was ; rank 1 reached at layer 62 (of 62).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 | 33 | 34 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 | 47 | 48 | 49 | 50 | 51 | 52 | 53 | 54 | 55 | 56 | 57 | 58 | 59 | 60 | 61 | 62 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 123163 | 248080 | 184778 | 208339 | 124945 | 128183 | 25324 | 76784 | 40361 | 177268 | 12238 | 79859 | 40915 | 115326 | 148736 | 146122 | 39591 | 98717 | 85457 | 60504 | 4090 | 3764 | 1406 | 1423 | 8526 | 36283 | 67338 | 92314 | 84992 | 95558 | 146109 | 161981 | 86650 | 196743 | 235260 | 220269 | 165829 | 222664 | 221410 | 242705 | 231603 | 244864 | 246271 | 231894 | 240700 | 206916 | 248095 | 248260 | 226853 | 247506 | 245124 | 244572 | 247578 | 239224 | 239569 | 230384 | 228065 | 223653 | 205285 | 73071 | 67615 | 26991 | 1 |