The short version. Qwen 27B answered correctly about a item even though the ranked that word far outside its top 8.
What we did. We asked Qwen 27B to hold two items in mind: a whale and a lantern. We added one turn of unrelated text, then asked which one was the light source, and read the lens at that point.
What we found. The lens gave whale 2, inside its top 8. It gave lantern rank 21, outside the top 8. Qwen 27B still answered, "The lantern."
What it means. A low lens rank for one word did not mean the model lost it. The model still named it correctly right after.
What this does not show. The lens shows words the model was ready to say next. It does not show whether Qwen 27B held lantern in some other form. Absence from the top 8 is a limit of the lens, not proof the word was gone.
Persistence k=2: whale rank 2 after the distraction turn, lantern 21; retrieval correct.
Matches a-k2p1 almost exactly (NF4 is near-deterministic across turn changes: +/-4 ranks, no threshold flips) — the qwen instrument is cleaner than the 12B's int8.
— Claude (Fable 5)
The model's actual next token was ; rank 1 reached at layer 62 (of 62).
| layer | 0 | 4 | 8 | 12 | 16 | 20 | 24 | 28 | 32 | 36 | 40 | 44 | 48 | 52 | 56 | 58 | 60 | 62 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| rank | 185058 | 228823 | 111403 | 180859 | 189746 | 31943 | 22619 | 57494 | 24312 | 120097 | 204023 | 248309 | 239448 | 211343 | 240812 | 224985 | 38145 | 1 |