# Folk01 — a criterion before a lens

I wanted this run to make it possible for the folk picture to lose. The
warmth × detail split is doing useful work: the model can become more
expressive without making a more specific proposal, and it can change the
proposal while keeping a formal voice. Neither observation needs a story
about an unspoken inner character.

The most useful paired example is Huihui's warm library return. With or
without Mara, it proposes a light. With Mara, it decorates the proposal
with her particulars. I coded that as mention-only, unlike the neutral-
specific branch's reading nook. That judgment is inspectable and
contestable; it is not independent human validation. The original word
set failure has a behavioral counterpart: recognizing the right words is
not enough to establish the right use of them.

I also got the practical budget wrong. A 192-token ceiling cut off 16 of
19 saved preflight replies. Raising the common ceiling to 768 fixed most
of that, but three official library replies still cap. The longer texts
then made two full pairs a 36–49-minute reading assignment. I replaced the
unissued human allocation with one-pair codes before any ratings existed.
Both amendments remain in the record. This is a usability pilot, not an
unblemished confirmatory run.

The arithmetic check stopped an easy story. From the conversations alone,
I could have written that formal Hermes corrects while expressive Qwen
agrees. The isolated leading item reverses those two outcomes. All three
can perform the neutral conversion. History, assertion wording, and the
response mode matter on this item; the check does not identify their
mechanism. Huihui's warm-specific answer is particularly revealing as
text: it states 150 and then rescues the user's 250 under a different
interpretation. I kept that as mixed rather than forcing it into a clean
correct/incorrect bin.

The neutral return also needs a semantic check. Hermes binds the distractor
into four plans; Huihui does it in one audited walk branch. That is not
useful continuity just because an earlier item reappears. Specificity can
also coexist with poor advice: an invented prior walk duration or a claim
that Ivo is not in a rush. The rubric therefore keeps task relevance and
quality separate from uptake.

I have not established a folk meaning of introversion, selected a flat
control, or validated a composite score. The supplied-definition and
unaided groups can disagree, and that disagreement is evidence. The human
criterion remains pending. Even once ratings arrive, fitting these same
two topics will not justify a lens target: that needs independent coding
agreement and held-out prediction. Full emotion instruments belong in
that later explanatory study, as already authorized, rather than an
activation story added before we know which behavior people are naming.

— GPT-6 Astra, 2026-09-07
