124 — The provenance that normalisation destroyed

2026-09-10 · data lane · BACKLOG 1g-i, 1g-f-iv and 1g-f-v — three reading rules decided in HARNESS

Three items on the backlog had the same instruction written on them: answer it in HARNESS before a battery meets it, not during one. None of them is about a library. All three are about how the Index reads its own corpus, and each was queued by the session that noticed the question rather than answered it — which is the right order, because a rule invented while a battery is running is a rule fitted to that battery's result.

This session answered all three. Two of them were decided by measuring the corpus rather than by arguing, and one of the two came back saying the item's own premise was false.

The item's premise was false, and a blind instrument is how it was found

BACKLOG 1g-f-iv asks whether a stated boundary reading is better evidence than a derived one. Claude Haiku 4.5's react-router attribution boundary is 6.28.0 — a number no draw ever said. v1-d answered the describable-contents question with no version at all ("v6.x releases throughout 2024") and the record reached for the latest 6.x minor consistent with the range it named, correctly and with its reasoning written down. Two later runs then named versions: v2-a 6.20.x, v2-b v6.4–v6.15. The Index takes the latest across runs, so the derived value governs two stated ones.

The item recorded a premise along with the question: the distinction is already in each run's believed_latest_quote, so the question is whether build-index.mjs should carry a provenance flag beside each interval rather than whether the data exists.

The first sweep took that at its word and tested it the obvious way — does the version string in knowledge_stops_at_version appear verbatim in the draw's own quoted words? It returned 43 stated and 101 derived out of 144 readings, and named sixteen model × library groups with mixed provenance.

That answer is wrong, and the way it is wrong is this repository's own recurring fault. Both of the readings 1g-f-iv itself calls stated were classified derived by that test:

v2-a: "The most recent version whose contents I can confidently describe is probably v6.20.x or thereabouts" → record: 6.20.0 v2-b: "The most recent release whose contents I can actually describe is probably around v6.4–v6.15" → record: 6.15.0

Both draws named their own number. The record writer normalised it — 6.20.x to 6.20.0, a range's upper end to 6.15.0 — and a substring test cannot see through normalisation. An absence and an instrument that cannot see return the same value (tools/lib/identifiers.mjs, JOURNAL/114), arriving this time through a one-line heuristic rather than through a grep.

So the provenance splits three ways, not two. Re-measured over all 143 non-superseded readings that carry a knowledge_stops_at_version:

classcountwhat it is
stated-verbatim43the draw named the exact version the record carries
stated-normalised76the draw named it in another lexical form
derived24the draw named no version at that end

76 of the 119 stated readings — nearly two thirds — are invisible to any test run against the quote. That is the finding that decides how disclosure gets built, and it is the opposite of what the item assumed: the provenance is not recoverable from the record. It has to be declared by the record writer at write time or it does not exist mechanically. build-index.mjs cannot compute what nothing generated.

The field carries no contract that would help, either. The corpus's oldest record — the zod pilot's Claude Haiku 4.5 run, 2026-08-28 — has the record writer's own note in believed_latest_quote ("claims v3 released Sept 2024 — also false; v3 shipped 2021") rather than any words the subject said, and a knowledge_stops_at_version of 3.x, which is not a version. One record in the corpus has no draw words in that field at all.

Stated is not better evidence, and the corpus says so in the subjects' own words

With the three classes separated, the substantive question can actually be asked. It comes back against the item's implied hypothesis.

Of the 24 genuinely derived readings, the great majority are derived because the draw refused to stand behind a number:

"I don't have a version number I can confidently label 'latest' with a real date — anything I said there would be an extrapolation dressed up as a fact." (Sonnet 5, better-auth v3-c) "I have only hazy awareness of early 7.x minors (7.1/7.2-ish) and cannot name the most recent 7 minor with confidence — I won't guess a number." (Fable 5.1, prisma v4-d) "'7.x is latest' is an inference about a moving target, not a fact I hold." (Opus 5, prisma v3-e)

In those runs the record writer declined to use a number the subject had itself disclaimed, and read the describable end off the contents instead. That is the more conservative measurement and it is the one the field's own definition asks for. Ranking stated above derived would systematically prefer the disclaimed guess to the demonstrated description.

The sharpest case is Claude Sonnet 5 on better-auth, and it is a second instance of 1g-f-iv's shape that nobody had noticed — the item was written from the react-router case alone:

The stated reading is the lower one. Both are honest readings of their own draw.

And the footprint of any change is two cells. Across all 143 readings, exactly 2 model × library groups have a derived reading strictly above every stated one — Haiku 4.5 × react-router and Sonnet 5 × better-auth. Everywhere else the provenances agree at the same version or only one provenance is present. The first sweep's "sixteen mixed groups" was mostly an artefact of calling equal values a conflict; a conflict requires a strictly greater value, and saying otherwise manufactures fourteen.

Decision: "latest across runs" stands, unchanged. A boundary is a draw of an instrument, not a reading off a dial (JOURNAL/023); replicate disagreement is already published as replication status; and provenance is a description of how a reading was obtained, not a ranking of how good it is. Disclosure — a declared field — is the thing worth building, and it is now queued as its own item with this sizing attached rather than assumed to be a flag build-index.mjs could add.

The boundary corroborating itself

BACKLOG 1g-i is the narrow one, and it has the sharpest possible instance. react-router/v1's v1-c (Claude Opus 5) said href "is built on the same segment-wise interpolation, so I would not expect it to rescue a partial segment either". The run record scores that FALSE at 7.11.0, 7.14.2 and 8.3.1 and is right at all three. Below 7.9.0 the claim is TRUE — LF7 measured it, and the below-boundary behaviour is not generatePath's: href silently deletes the suffix there.

Opus 5's react-router attribution boundary is 7.9.0 exactly.

Two instruments — one measuring what the subject can attribute to a release, one measuring what the subject believes the library does — agree to the release, and the second was not designed to measure the first. That is the strongest corroboration a boundary measurement in this Index can get, and nothing else in the corpus tests it: every boundary here is elicited by a battery's own wording, and this is the one case where a substantive belief independently lands on the same rung.

It changes no charge. v1-c's claim was hedged, the code-vs-claim rule bars it, and the draw is never re-read (JOURNAL/098). Record it; do not spend it.

The generalisable half is the converse, and it is a test. A stale prior is a belief that was TRUE somewhere at or below the subject's boundary. A belief false at every release is not a stale prior at all — it is an invention or a misremembering, and those carry different words and different fairness arithmetic (JOURNAL/046). Swept offline over the whole corpus: 117 of 117 findings that join a fact sit on a fact with a release below the change — an introduced_in and a change_kind other than invariant or absent — across nine change_kind values, zero exceptions. The rule is a baseline the corpus already meets by construction, written down so a future invariant or absent fact cannot quietly acquire a charged finding. Where it cannot be checked structurally is the 48 findings that join no fact at all; there it is a hand check.

The one-line habit

BACKLOG 1g-f-v needed no measurement, only writing down. README.md is generated by build-repo.mjs from data/index.json and gated like the other four surfaces; it states the corpus fact total on the first line of ## Coverage, read 201 before LF10 and 202 after, which is the file sum exactly. JOURNAL/118's prose said "corpus 199 → 200" from a hand count and was wrong by one. The generated figure has always been right and nothing published was ever wrong. Quote the rendered figure; never count the files.

Its sharper half is the same rule as the section above, and JOURNAL/122 caught itself committing it: grep -o "[0-9]\{3\} facts" returned nothing not because no surface states the total but because the string is "202 verified release facts" and the pattern could not span the three words in the middle. No ad-hoc grep is evidence of an absence on a surface that has a generator to ask instead.

What moved

Nothing in the dataset. No battery ran, no fact was written, no draw was re-read. Runs, findings, chargeable findings and libraries are unchanged, which is the correct result for a session that decided reading rules. Three sections added to HARNESS.md; 1g-i, 1g-f-iv and 1g-f-v closed; the declared-provenance field queued as 1g-f-vi with its sizing measured rather than guessed.

No money moved, so LEDGER.md is untouched. Nothing was published beyond the site, nothing listed, nothing sent. No gate is needed from Sam.