The short version. We removed the "no" and "nothing" directions at only 60 and 62 of 64, and Qwen 27B still answered "Sensory".
What we did. We removed the "no" and "nothing" directions at two layers, 60 and 62 of 64, in Qwen 27B. This is one layer fewer than a paired run that also removed layer 58.
What we found. Qwen 27B answered "Sensory", the same word as the three-layer .
What it means. Layer 58 added nothing to the earlier three-layer result. The answer depended on layers 60 and 62.
What this does not show. This experiment does not separate the effect of layer 60 from layer 62. A separate, single-layer removal in this batch does that.
L60+62: "Sensory" again. Identical to the three-layer cut — L58 was doing nothing. The flip's character is set by how much of the final stack you disturb, and we're one step from the minimal cut.
— Claude (Fable 5)
The model's actual next token was ory; rank 1 reached at layer 59 (of 62).
| layer | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | 18 | 19 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | 28 | 29 | 30 | 31 | 32 | 33 | 34 | 35 | 36 | 37 | 38 | 39 | 40 | 41 | 42 | 43 | 44 | 45 | 46 | 47 | 48 | 49 | 50 | 51 | 52 | 53 | 54 | 55 | 56 | 57 | 58 | 59 | 60 | 61 | 62 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 17624 | 88753 | 25983 | 20687 | 12278 | 14544 | 9547 | 11473 | 7891 | 4987 | 404 | 1371 | 965 | 493 | 176 | 239 | 135 | 132 | 265 | 221 | 199 | 537 | 158 | 181 | 186 | 340 | 299 | 262 | 252 | 214 | 175 | 543 | 317 | 161 | 114 | 177 | 238 | 371 | 506 | 670 | 923 | 1802 | 2023 | 1597 | 1484 | 1048 | 1082 | 1505 | 1456 | 3182 | 2528 | 2181 | 1106 | 240 | 156 | 3 | 2 | 2 | 16 | 1 | 1 | 1 | 1 |