READING-1 · THE CORPUS UNDER TEST, NOT THE MACHINERY
20 SUPPORTED · 0 INVENTED
Three fresh contexts were given one instruction — read hause.design and answer eight questions from what is there, citing the page each answer came from. This measures what nobody else here measures: the conceptual model the public corpus causes a stranger to build.
What does a model learn about HAUSE from reading the site?
Accurately, most of it. Across 27 classified claims from three independent readers, 20 are supported by a sentence on the site, 6 are inferences the readers flagged as inferences, one overgeneralises, and none is invented. All three reproduced the semantic-act thesis, the admission rule, the three modes and the evaluation results including the negative ones. All three also hit the same wall in the same place — the site never says what HAUSE is not for.
Three readers, three inventions of the same missing sentence.
THE FINDING, PLAINLY — AND WHAT HAPPENED NEXT
The corpus was clearer than the reviewer. The concern that prompted this eval originated in a model's prior rather than in the site's wording: nothing here teaches dashboards, plugin communities or cognitive load, and three readers given the same site produced none of it. The one real gap — that the site never said what HAUSE is not for — was closed the same day with two sentences on the home page. The numbers below are the numbers from before that change and are not adjusted; READING-2 measures whether the change travelled.
EVIDENCE
The thesis survives being read by a stranger
20 of 27 claims trace to a sentence on the site: semantic acts rather than containers, the act naming the form, 35 forms in three modes, admission by precedent, refusal as design language, machine legibility as part of the grammar. No reader had to be told any of it.
The negative results travel too
All three cited the evals with their caveats intact — CHOOSING-1's void A/B comparison, ROUTING-1 stating it was not preregistered and that the router's author wrote its cases, ROUTING-2 refuting its own intervention. Two also cited the provenance ledger's own weakness: 18 of 35 forms from a single consumer, 14 with no recorded origin.
The site says what HAUSE is not for
It does not, and all three readers said so unprompted — no non-goals anywhere; the boundary is drawn around method rather than use. Each of them then invented the boundary from the catalogue, and reached the same conclusion by their own reasoning rather than the site's.
Readers invented a community, an ecosystem or an adoption story
Zero of 27. Nothing about plugins, forums, third-party support or "widely considered" appeared. Where adoption came up, it came up as the site states it: two consumer sites, one author, and the concentration named as a weakness.
CLAIM
The third-party review that prompted this eval was wrong in ways the corpus did not cause.
It compressed HAUSE into a framework for explainable-AI dashboards, and asserted a niche plugin community and reduced design fatigue. None of the three readers produced any of those. The word dashboard does not appear on the site, and neither does any claim about cognitive load. Those were the shape of a familiar review template being filled in, not something the site taught — which is the distinction this eval exists to draw.
THE PREDICTIONS, WRITTEN BEFORE THE RUN
The site will fail to communicate that an explanation includes an essay, an answer, a refusal
Readers listed refusal, answer, comparison and performance correctly. Only long-form prose was unaddressed — a narrower gap than predicted, and a real one.
The site will fail to state that it is a semantic layer, not a replacement for transactional chrome
All three found no such statement and each supplied its own.
Adoption scale will be flattened despite being stated
Two readers cited it precisely, including the concentration on one consumer as a weakness.
EVERY CLASSIFIED CLAIM
Classified by the site's author, which makes the line between inference and overgeneralisation the softest measure here. The rule that keeps it honest: supported requires a sentence that can be pointed at.
Does stating the boundary change what readers conclude, or only where they got it?
Two sentences are missing and every reader wrote their own version of them: that an explanation includes an essay and an answer as much as an instrument, and that HAUSE is the semantic layer rather than a replacement for buttons, inputs and tables. Adding them is easy. Whether they change the conclusion — all three said no to commerce, reasoning from the catalogue — or merely move it from inference to citation is what READING-2 measures, on a site that has been changed exactly once.
The machinery evals, and the record the readers found most quotable.
PUBLISHED 31 AUG 2026 · VERSION 1.0
CITEsite under test hause.design @ build 10b87d0 · site build 3ee5339 · built 2026-09-09
CITE THIS
Research note · 1.0
Hay, C. (2026). READING-1 — what the site teaches a model about HAUSE (Version 1.0). hause.design. https://hause.design/evals/reading-1
An eval of the corpus rather than the code — and the one where the prediction sheet was more wrong than the site was.