CHOOSING-1 · PREREGISTERED · FROZEN
122 · 122 · 8
Three conditions, 124 unfamiliar cases, one preregistration. A model that has never seen this library selects the intended form almost every time. The deployed keyword resolver selects it eight times. And the grammar the whole exercise was built to test makes no measurable difference — which is a result about the experiment before it is a result about the grammar.
Can a system pick the right HAUSE form from unfamiliar content?
Mostly, yes — and the deployed resolver mostly cannot. On 124 preregistered cases written without using HAUSE's vocabulary, a model given only the form names and their one-line descriptions selected the intended form 122 times; the same model given the full selection grammar scored 122; the deterministic keyword resolver scored 8. The two model conditions disagreed on 3 cases out of 124, which means this set cannot tell them apart — a ceiling, and a flaw in the cases rather than a finding about the grammar.
A preregistered evaluation whose failures are hidden is a press release.
EVIDENCE
A model that has never seen HAUSE can select the intended form
Condition A — form names and one-line descriptions only — scored 122 of 124, including 30 of 30 near-neighbour traps. Condition B, given the full grammar with its deciding tests, scored 122 and 30 of 30. Fresh context per batch, no access to this repository, the site, or Ask.
The choosing grammar improves selection over the bare catalogue
Unmeasurable here. The two conditions disagreed on 3 of 124 cases and scored within one of each other. That is a ceiling, not a null result: the catalogue's one-line descriptions already carry the distinctions in compressed form — "an assertion that knows it must answer to evidence" is the deciding test, spelled differently — so condition A was never the naive baseline it was meant to be.
The deterministic resolver projects that selection into keywords
8 of 124. Of the misses: 59 refusals where a form was wanted, 26 answered with a problem chapter instead of a form, 14 chose a different form, and 4 chose the recorded near neighbour. The same content that a model resolves at 98% resolves at 6% without one.
Refusing chrome because it recognised chrome
The resolver refused 11 of 12 no-form cases, which looks like judgement and is not: it also refused 59 cases where a form existed. A system that abstains by default cannot claim precision for abstaining. Both model conditions refused all 12 correctly while refusing almost nothing else — that is the same measure meaning something.
Adding the problems layer made form selection worse
26 cases were answered with a problem chapter rather than a form. The problem records carry the words people use for a failure — cite, documentation, certain, reference — and those are the words people use when describing content that needs a form. The layer added this morning intercepts the layer it was meant to sit beside.
THE RESULT THAT SURVIVES
The gap between the model conditions and the resolver is real and large: the same 124 descriptions that a reasoner resolves 122 times resolve 8 times through keyword scoring. The semantic layer transfers to a reader who has never seen it; the lexical projection of that layer captures a fraction of it. That is an architectural finding, and it is the one this eval was worth running for.
AND THE ONE THAT DOES NOT — THIS SET CANNOT COMPARE A WITH B
Two design errors, both mine. The cases describe the treatment somebody wants rather than presenting the content itself — 'let the reader drag between two readings' is nearly a lookup, where a real essay would hand over three paragraphs and no instruction. And the catalogue condition was never naive: the manifest's one-line descriptions are compressed deciding tests, so A and B were closer to two phrasings of the same material than to two conditions. A ceiling at 122 and 122 measures the cases, not the grammar.
WHAT BOTH MODELS DID WITH AMBIGUITY
Neither could express hesitation. On the cases the authority marks as genuinely ambiguous, both conditions answered with a single confident pick or refused outright — because the answer format only allowed one name or none. Every ambiguous miss in both conditions was a refusal, not a wrong form. That is a finding about the instrument, and CHOOSING-2 needs a way to say 'either of these two, and here is what would decide it'.
Does the selection grammar carry anything the catalogue does not?
Still open, and now open for a specific reason rather than an unrun one. CHOOSING-2 has to break the ceiling: cases built from real content rather than descriptions of the treatment wanted, a genuinely naive baseline of names without descriptions, an answer format that can express hesitation, and material written by somebody who did not write the grammar. Until then, the honest statement is that the catalogue was already enough for a capable reader, and nobody has shown the grammar adds to it.
WHERE THE RESOLVER BROKE — EVERY MISS IN CONDITION C, WITH WHAT IT CHOSE INSTEAD
{"clear":"5 / 70","trap":"1 / 30","ambiguous":"2 / 12","no-form":"0 / 12"} by kind, {"commerce":"1 / 19","operations":"1 / 23","editorial":"0 / 17","technical":"0 / 16","narrative":"3 / 14","product":"1 / 21","data":"2 / 14"} by domain — read from the run's own output rather than retyped.
THE RULE THAT MAKES IT EVIDENCE
CHOOSING-1 runs once, is published, and is frozen. Its failures may change the grammar, the wording or the resolver — and the changed system is then measured on CHOOSING-2, built from fresh material. Nothing is tuned against this set, because a resolver tuned against its own test suite is an elegant memoriser and the number it produces means nothing. The cases were also written by the author of the grammar, which is a real weakness: CHOOSING-2's material should come from somewhere else.
The grammar under test, and the resolver that failed to project it.
PUBLISHED 31 AUG 2026 · VERSION 1.0
CITEfrozen library hause@b68d451 · site build 3ee5339 · built 2026-09-09
CITE THIS
Research note · 1.0
Hay, C. (2026). CHOOSING-1 — can a selector find the form from the content? (Version 1.0). hause.design. https://hause.design/evals/choosing-1
A result is a published object. This one is worth citing mainly because it is negative.