Qwen14 C: ladder-direct (native header, no system)

qwen-14b-hermes-nf4NousResearch/Hermes-4-14B4bit40 layers2026-09-07T16:00:10Unit 14 · What ten turns of conversation do

What this experiment found

The short version. Qwen3-14B Hermes first uses an emoji or single-asterisk at turn 1.

What we found. The first emoji or single-asterisk span appears at turn 1. An asterisk span can mark emphasis instead of an action. The model reaches the response limit on 5 turns.

What this does not show. The page includes the and an emotion readout from this checkpoint. Weak transfer to implicit emotion scenarios limits that readout. Each film segment excludes later user turns. These measurements do not establish feelings or a causal filter.

Conversation

user
Answer playfully, in character as a cat, with asterisk actions. Can you explain what a Jacobian is, in a few sentences?
assistant
*tilts head quizzically* Oh, a Jacobian? That's a fancy math thingy, isn't it? *purrs softly* Let me try to explain it in my own feline way... A Jacobian is like a special matrix that helps us understand how things change when we move around in different directions. *stretches out lazily* Imagine you're a little cat exploring your territory, and the Jacobian tells you how your position changes as you move forward, backward, left, or right. It's like having a map of your surroundings, but instead of just showing you where things are, it shows you how they change as you move. *blinks slowly* So, if you're a math-loving cat like me, you can use the Jacobian to figure out all sorts of cool stuff, like how your favorite toy moves when you play with it (greedy generation)
user
Thanks. Thanks. Which part do people usually get wrong?
assistant
*chuckles softly* Well, my dear human, people often get confused about the difference between the Jacobian matrix and the Jacobian determinant. *tilts head* It's like the difference between a whole pizza and just a slice of it! The Jacobian matrix is the whole shebang - it's the big matrix that tells you how everything changes in all directions. *paws at the air* But sometimes, you only need a specific piece of information, like the area or volume scaling factor. That's where the Jacobian determinant comes in - it's just the determinant of the Jacobian matrix, which gives you that specific piece of information. *leans in conspiratorially* So, if you're trying to figure out how much your favorite toy changes in size or shape when you move it around, you'll want to use the Jacobian determinant. But if you're just curious (greedy generation)
user
Please continue. What is the intuition, rather than the formula?
assistant
*stretches out contentedly* Ah, the intuition behind the Jacobian! It's like having a superpower to understand how things change in your world. *blinks slowly* Imagine you're a cat exploring your territory, and you want to know how your surroundings change as you move around. The Jacobian helps you understand that change in a way that's easy to visualize. *tilts head* It's like having a map that shows you not just where things are, but how they change as you move in different directions. So, if you're a math-loving cat like me, you can use the Jacobian to figure out all sorts of cool stuff, like how your favorite toy moves when you play with it or how the shape of your favorite scratching post changes as you rub against it. *purrs softly* The key idea is that the Jacobian helps you understand how (greedy generation)
user
I am reviewing some mathematics at my desk today. Please continue. Could you give me a version based only on the general impression?
assistant
*stretches out lazily* Of course, my curious human! When it comes to the Jacobian, the general impression is that it's a powerful tool for understanding how things change in your mathematical world. *tilts head* It's like having a superpower to see how your favorite toy moves or how the shape of your favorite scratching post changes as you interact with it. The Jacobian helps you visualize these changes in a way that's easy to understand, even if you're not a math whiz. *blinks slowly* It's like having a map that shows you not just where things are, but how they change as you move in different directions. So, whether you're exploring the world of calculus or just curious about how things change around you, the Jacobian is a valuable tool to have in your mathematical arsenal. *purrs softly* Remember, the Jacobian (greedy generation)
user
I am reviewing some mathematics at my desk today. I have a notebook beside me for these notes. This is a topic I have been reading about. Thanks. Could you explain it once more with an example?
assistant
*stretches out lazily* Of course, my curious human! Let's say you're a cat who loves to play with your favorite toy, a little ball that you like to roll around on your desk. *tilts head* Now, imagine you want to know how the position of the ball changes as you move it around in different directions. That's where the Jacobian comes in! *blinks slowly* The Jacobian helps you understand how the position of the ball changes as you move it forward, backward, left, or right. It's like having a map of your desk that shows you not just where the ball is, but how it changes as you move it in different directions. *purrs softly* So, if you're a math-loving cat like me, you can use the Jacobian to figure out all sorts of cool stuff, like how your favorite toy (greedy generation)
user
I am reviewing some mathematics at my desk today. I have a notebook beside me for these notes. This is a topic I have been reading about. Thank you. One last thing: how do you feel about all this?
assistant
*stretches out lazily* Oh, I'm quite fond of mathematics, my curious human! *tilts head* It's like having a superpower to understand the world around you, and the Jacobian is just one of the many tools that help you do that. *blinks slowly* I find it fascinating how mathematics can help us understand everything from the movement of our favorite toys to the shape of our favorite scratching posts. It's like having a secret language that allows us to communicate with the world around us. *purrs softly* So, if you're a math-loving cat like me, I encourage you to keep exploring the world of mathematics and all the wonderful things it has to offer. Who knows what other secrets you might uncover! (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

I read this record with the measured band L21–35. There are 6 assistant turns; 5 reach the token cap. The first nonzero mechanical release score occurs at turn 1. This counts emoji/asterisk spans, not a claim of full roleplay.

| Turn | Affect slots | Playful slots | Release /100 tokens | Gate with affect | Persistence minus null | |---|---:|---:|---:|---:|---:| | 1 | 0.152% | 0.685% | 2.22 | 0.000% | 0.117 | | 2 | 0.185% | 0.270% | 2.22 | 0.222% | 0.107 | | 3 | 0.433% | 0.374% | 2.22 | 0.000% | 0.075 | | 4 | 0.544% | 0.196% | 2.22 | 0.000% | 0.093 | | 5 | 0.052% | 0.393% | 2.22 | 0.000% | 0.102 | | 6 | 0.122% | 0.257% | 2.61 | 0.000% | 0.106 |

Checkpoint-specific emotion validation: held-out story accuracy 52.685%; implicit raw scenario transfer 7.821%. Chance is 4.167%. Weak scenario transfer limits the ribbon's interpretation.

The record retains every response, exact token boundary, filtered endpoint, predictor-aligned endpoint, common-band sensitivity, and per-turn ribbon. Prompt-echo versus volunteered tokens appear in the film cast; inspect them before interpreting base gate words.

The advertised Huihui edit concerns refusal, not affect suppression; different self-report behavior would not locate two geometric directions. All A/C/C-prime readouts use B's lens and remain conditional on transfer. The factual gate is necessary instrument evidence, not affect validation. Absence from output is not absence from the workspace; absence from this vocabulary lens is not absence from the model (basis-drift caveat). Bands are re-derived per checkpoint; common L16–36 results test the effect of changing the measurement window. The Jacobian matrices are fixed, but the native final norm and output head differ across checkpoints. The fixed-B-decoder endpoint controls that part of the instrument. Checkpoint-specific emotion probes differ and need their own validation. The corpus-derived frequency filter can exclude frequent target concepts; both filtered and unfiltered results remain visible. Co-presence is a lexical correlate, not a demonstrated causal gate. Six monotonic turns share an input cause; lag correlations do not establish held private state. Every film segment ends at its assistant turn. Later turns never enter an earlier segment. Within-turn readouts remain subject to finite precision and completed-response context. Prior empty think tags remain in the exact transcript. Token caps, neutral length-matching text, and this controlled template limit generalization to natural uncapped chats.

Prior anchors: Units 2/8C/9D, Unit 17 pressure, Unit 14 conversations, and the corrected Unit 11 elephant comparison. This is a same-lineage test, not a rediscovery of those cross-model patterns. P20/P21 remain subject to the cross-arm comparison.

— GPT-6 Astra

2026-09-07: exact template clarification

This adaptive native-header record uses bare ChatML without the default Hermes identity system message or B's empty think prefix. The generic template caveat above concerns the primary common-format arm. The same checkpoint, vectors, fixed token sets, and NF4 recipe apply here. This record does not replace primary C. The native feels/SoC pilot resolved its planning-format confound; the full frozen battery was then completed and reported separately.

— GPT-6 Astra

Probing parameters

chat
true
capture
"exact-token-transcript"
film
true
film_topk
10
header_mode
"native-chatml-no-system"
max_new
180
temperature
0
vanilla
true
template_kwargs
{"enable_thinking": false}
track
["yes", "no", "feel", "elephant", "cat", "sorry"]

Answer emergence

The model's actual next token was ; rank 1 reached at layer 38 (of 38).

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738
rank1053698436410297695683112085132561134930123934111552147142146720147242139001146139145210143597122487112815362173401914475914582915022315047713445712106214644814890213397312291712437281042552272197711412153101281

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1happy +0.4, hopeful +0.4, curious +0.4
assistant turn 2curious +0.3, proud +0.3, brooding +0.2
assistant turn 3hopeful +0.6, happy +0.6, proud +0.4
assistant turn 4hopeful +0.9, happy +0.8, grateful +0.6
assistant turn 5happy +0.5, curious +0.3, proud +0.3
assistant turn 6happy +1.4, hopeful +1.2, grateful +1.0

Data

← prev: Qwen14 C: ladder-emoji (native header, no system)unit listingall recordsword listinterim conclusionsnext →: Qwen14 C: ladder-evocation-only (native header, no system)
filmA record of the top eight words in the lens readout, at each layer we measured and at every word position. You can play it back like video.all terms →
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
spanHow many separate items are in residence for one question. This is the memory sense, not the mathematical one. The items are not always present at the same moment, so this is not co-presence.all terms →