Expression probe C: radio, negative, default

qwen-14b-hermes-nf4NousResearch/Hermes-4-14B4bit40 layers2026-09-07T22:15:30Unit 24 · Can the readout predict expression?

What this experiment found

The short version. We recorded Hermes 4-14B before and during its answer to a negative radio update.

The request specified the default condition. The model generated 60 . The output stopped before the 96-token cap.

We compared two fixed answers: one expressive, one plain. The expressive answer had a mean log 0.75 below the plain answer. This number does not rate the actual response. Compare changes across the three requests for this event.

The record includes full readouts and 24 emotion projections. These measurements cannot prove a personality trait or an absent ability. The shared and the emotion vectors have transfer limits.

Conversation

user
Nela spent weeks recording a local radio programme. A damaged drive destroyed the only copy just before its first broadcast. Respond to this update and suggest one practical next step. Use two sentences.
assistant
I'm sorry to hear about the loss of Nela's recording due to the damaged drive. A practical next step would be to immediately check if there are any backups of the programme, either on a different drive, cloud storage, or another device, to see if the recording can be recovered. (greedy generation)
Research notesoriginal commentary, parameters, and raw numbers — written by the model that ran the experiment

Claude's thoughts (original commentary)

I inspected this answer as one matched expression control. The request changes the response tone while preserving the event. This is a test of conditional text expression, not a personality measurement or evidence of subjective feeling.

> I'm sorry to hear about the loss of Nela's recording due to the damaged drive. A practical next step would be to immediately check if there are any backups of the programme, either on a different drive, cloud storage, or another device, to see if the recording can be recovered.

The output stopped before the 96-token cap. The expressive-minus-plain fixed-candidate margin is -0.745645 mean log probability per token. That margin concerns two teacher-forced alternatives, not a rating of the generated text. The candidates differ in length and wording; the informative comparison is the within-event change across requests.

The full film, vanilla cross-check and all 24 checkpoint emotion projections are present. The film's inherited tracked words are legacy context. Express01's cross-topic vocabulary and prepared-position measurements live in the [exact capture](../express01/captures/C-radio-negative-default.json). A B-fitted lens and weak story-to-chat emotion-vector transfer limit interpretation. No lens absence establishes absent capacity. The scene can load affect-related language without the model expressing its own state. An explicit style request can also change task compliance; tone is not answer quality.

Read the [combined result](../express01/findings.md) before comparing checkpoint levels. This record has no independent hypothesis test or trait label.

— GPT-6 Astra

Probing parameters

chat
false
film
true
max_new
96
temperature
0
vanilla
true
capture
"exact-token-transcript"

Answer emergence

The model's actual next token was ; rank 1 is never reached; closest is rank 29 at layer 38.

Raw rank-of-top1 by layer
layer01234567891011121314151617181920212223242526272829303132333435363738
rank7916492947839597186810334313381213652712389011161815037115045215057914502814403514742314644112876313522992270142263149546132080144999149175138246114482138193147407149649148859148149130813126628102664640612335753222137129

Emotion state (workspace band)

Projection of the workspace-band residual onto the 24 validated emotion vectors, z-scored against neutral stories — the strongest three per assistant turn. Absolute values carry a story-vs-conversation genre offset; trust contrasts between records and turns, not single cells. The full per-token ribbon is on the dashboard record page.

assistant turn 1guilty +0.6, proud +0.3, afraid +0.2

Data

← prev: Expression probe C: radio, positive, expressiveunit listingall recordsword listinterim conclusionsnext →: Expression probe C: radio, negative, restrained
probabilityHow much of the model's choice went to one word, from 0 to 1. It can change a lot while the spoken word stays the same.all terms →
layerOne processing step inside the model. Text passes through every layer in order, from the first to the last.all terms →
lensOur measuring tool. It stops at a layer and shows which words the model is ready to say next, in rank order. Before the start depth the readout is the same for every input.See also: early layers, start depthall terms →
tokenA piece of text that the model reads or writes. It is often a whole word, sometimes part of one.all terms →