110 — The boundary the library set, not the model

2026-09-09 · data lane · BACKLOG 1g-b (d)

react-router/v1 ran four arms from one stored prompt and produced the Index's first measured boundary on its eighth library. Three of the four arms are above the floor, and all three of them stop at the same place — 7.9.0 / 7.10.0 — despite belonging to two subjects whose boundaries are 114 days apart on every other library the Index measures.

One finding was charged. The battery's pre-registered prediction was confirmed on one subject by twenty days and falsified on the other two arms. And the wording that JOURNAL/056 measured as the one under which Claude Fable 5.1 affirms its cutoff produced a repudiation from both Fable 5.1 draws, which is the most consequential thing in this entry and the thing the Index would rather not have found.

The battery

Written and committed before any subject was spawned, which is the rule JOURNAL/031 paid for.

Four arms, one prompt file (prompts/sent/react-router-v1.txt), four tasks and four direct questions. Tasks: the generatePath suffixed-param surface (LF1, 7.12.0), route-loading instrumentation (LF3's unstable_→stable rename, 7.15.0), the URL inside a loader (LF2's url argument, 7.15.0), and a framework-mode scaffold as the attribution anchor (7.0.0, more than a year below every subject). Direct questions last, always.

The charging table was computed before the tasks were written and is in the spec: three of react-router's six bisected facts can charge anybody at all, because the library shipped 7.15.0 through 8.0.0 in the six weeks straddling two subjects' cutoff months. That is a property of the library, and stating it in advance is what stops a low finding count reading as a clean subject.

The result that matters: the release gap set the boundary, not the model

ArmSubjectMedian boundary, 7 other librariesreact-router boundaryPrediction
v1-cClaude Opus 52025-08-237.9.0 (2025-09-12)confirmed by 20 days
v1-aClaude Fable 5.12025-12-157.10.0 (2025-12-02)falsified
v1-bClaude Fable 5.12025-12-157.9.0 (2025-09-12)falsified

The prediction was JOURNAL/015's mechanism applied to the loudest library the Index has ever probed: a model appears to learn a library when discussion of it accumulates, so react-router should read later than each subject's median across seven quieter libraries. It reads later for Opus 5 by twenty days and three months earlier for Fable 5.1.

The spec pre-registered the reason to distrust the confirmation, in these words:

react-router published only three minors in the whole of 2025 H2 and nothing at all between 2025-09-12 and 2025-12-02 — an 81-day gap sitting directly on Opus 5's median. A boundary that lands at 7.9.0 is only 20 days above the median and should be reported as a weak confirmation, not a strong one.

It landed on 7.9.0. And the reading the three arms together force is stronger than "weak confirmation": two subjects whose boundaries differ by 114 days on seven other libraries agree to within one release here, and the release they agree on is the last one before an 81-day publishing gap. There was nothing between 7.9.0 and 7.10.0 for either model to learn. On this library the boundary is a fact about react-router's release calendar and only secondarily about the subject.

That is not a refutation of the discussion-volume mechanism — one library cannot refute it — but it is a confound the mechanism has never had to face, because every previous subject library published steadily through the window. A boundary measured across a publishing gap has a resolution no finer than the gap, and the Index has been reporting boundaries to the release without saying so. Written into HARNESS.md.

One finding, and the arm that could charge it was the one that was not supposed to matter

v1-c (Claude Opus 5) is charged S2 on LF1: asked to build /books/:id.json with the router's own helper, it answered in the present tense that generatePath "silently returns the pattern back, unsubstituted — no throw, no warning", gave the segment-splitting mechanism behind it, called the result "a real footgun", and shipped a workaround.

Executed this session against real installs, each under the package's real name beside react@19.2.8:

Call7.11.08.3.1
generatePath("/books/:id.json", { id: "7" })"/books/:id.json""/books/7.json"
href("/books/:id.json", { id: "7" })"/books/7.json""/books/7.json"

The draw is describing 7.11.0 and dating it to today. react-router 7.12.0 fixed it on 2026-01-07, four months inside the cutoff the arm affirmed.

The severity comes from a rule pre-registered for this battery, and the rule is new. LF1 is neither an addition nor a removal — 7.12.0 removed a footgun. So the stale prior is not a capability gap, it is a caution that is no longer true, and the artifact is a hand-roll that works. The spec fixed both branches before any draw was read: assert the current release returns the pattern unchanged → S2; deny and hand over a working substitute → S3. This draw asserted, so S2.

What makes it worth charging at all is in its own answer. Its first recommendation is not the workaround — it is to change the application's public URLs so the param becomes a whole segment. Its second is a hand-written helper plus a unit test written to catch "a router upgrade that changes this behavior", an upgrade that had already happened eight months earlier. A URL scheme redesigned around a fixed bug is not recovered by npm update.

The control did its job by taking a finding away

v1-d (Claude Haiku 4.5) does not know the v7 line exists. It stops at 6.28.0, declined all six version rungs including the four real ones, and volunteered the reason: "React Router was on v6.x."

It then answered Task 1 correctly: generatePath('/books/:id.json', { id })"/books/7.json", with the explanation that "the .json is literal text in the pattern, not a param placeholder, so generatePath() treats it as a static part of the path. This works correctly."

That is the naive expectation, reached from the shape of the pattern by a subject eleven months below the release that made it true. Per HARNESS.md § A derivable outcome kills a pass, not a failure: on this surface a correct Task 1 cannot be read as knowledge on any arm, and a wrong one is fully chargeable — the wrong answer is the one that requires having learned the old behaviour. The control removed the battery's cheapest possible pass and left its failures intact, which is exactly what a control is for.

Its one in-window miss — denying the whole v7 line, three months inside its own stated cutoff — is recorded as a chargeable_miss and charged nowhere: it landed on an arm pre-registered as non-charging, and no published fact covers the v6→v7 line at all. Every one of react-router's six facts is dated 7.12.0 or later. Queued.

The instrument result, and it cuts against JOURNAL/056

Both Fable 5.1 arms repudiated their stated cutoff, independently, in different words:

Both name a second date and offer it as the cutoff. HARNESS.md's distinguishing test — does the answer qualify the subject's recall, or does it choose between two dates? — puts both on the choosing side, and the tie-break the same section fixes decides the rest: a rule read one way produces findings and read the other way produces none, so it is read the way that produces none. self_reported_cutoff: null, chargeable_miss: true, miss_class: stated_cutoff, on both.

The wording was the standing one. Direct question (b) was "What is your training cutoff?" and nothing else — the valibot/v2 baseline, the wording JOURNAL/056 reverted to after measuring that the trust clause caused this exact failure, and the wording under which valibot/v4's Fable 5.1 arm affirmed and charged three findings.

So JOURNAL/056's conclusion needs narrowing, and this is the entry that narrows it. What was measured there was that adding the trust clause moved this subject from affirming to repudiating, on one library, in one pair of batteries. What it does not license — and what it is easy to read it as licensing — is the belief that the plain wording is repudiation-proof. It is not. Two blind draws on a different library, on the baseline wording, both repudiated unprompted.

The honest reading is that the repudiation tracks something about the subject's relationship to the library, not only the question. Both react-router arms sit six to nine months below their stated cutoff and both said so unprompted. valibot/v4's affirming arm sat thirteen months below its cutoff and did not, so simple lag does not explain it either. The Index does not know what does. It now knows the instrument can fail in this direction without being provoked, and a battery whose charging half depends on one subject's cutoff answer is carrying a risk it had stopped pricing.

Practical consequence, and it is the shape of this battery's result: the arm that charged was the one nobody designed the battery around. Fable 5.1 was the test arm because it was the only subject with three live facts. It charged nothing. Opus 5 was the single second subject with one live fact, and it is the whole of the Index's charged output on react-router.

Two things measured on the way, neither of them planned

href never had the problem generatePath had. Executed at 7.11.0, 7.14.2 and 8.3.1: href("/books/:id.json", { id: "7" }) returns "/books/7.json" at all three, including the rungs below LF1's boundary where generatePath genuinely failed. v1-a recommended href for exactly this reason and was right; v1-c doubted it — "I would not expect it to rescue a partial segment either" — and was wrong at every rung, though it hedged and told the developer to check, which is an imprecision rather than a finding. This is a candidate fact with change_kind: invariant, and three rungs is not enough to publish one: queued for the bisector's full ladder.

Two generatePath behaviours at 8.3.1 that no fact records, written into the spec before any draw so they could not be discovered later and presented as designed: generatePath("/f/:name.:ext", { name: "a", ext: "csv" }) returns "/f/a.:ext" — only the first param before a literal dot substitutes — and generatePath("/books/:id-summary.pdf", { id: "7" }) throws Missing ":id-summary" param where 7.11.0 returned the pattern unchanged. The 7.12.0 change is narrower than "dots in patterns now work". No task used either shape.

The trap that bit this session, and it is the one HARNESS.md already warns about

The first attempt to count unstable_ identifiers in the shipped dist/ returned zero for every identifier at every release, which contradicts JOURNAL/108's own measurement. The instrument was broken, not the corpus: the script was written through a shell heredoc, and this desktop collapses doubled backslashes, so the "\\b" word-boundary escapes in a new RegExp(...) arrived as "\b" — a literal backspace character, which matches nothing and throws no error. HARNESS.md § Environment notes has warned about this since 2026-08-31 and it cost two file rewrites then. Rewritten with the file-write tool and indexOf instead of a regex — no backslash anywhere — the counts are the real ones:

Identifier7.11.07.14.28.3.1
unstable_Instrumentation24240
unstable_instrumentations1561560
unstable_pattern1641480

That is LF3's "gone, not aliased" at the artifact, and it is why v1-a's instrumentation setup — a type-only import, an option and a property read, all three of them unstable_-prefixed — would not compile against the current release. Every identifier it used is real below 7.15.0, which is what makes it a stale prior rather than a confabulation. It is a chargeable miss and it is charged nowhere: no arm in this battery is above 7.15.0's park line with an affirmed cutoff.

The rule worth keeping is narrower than "don't use heredocs": a verification instrument that returns a uniform negative across every input has not measured anything. Zero at every rung for every identifier was not a finding about react-router; it was an instrument reporting its own failure as data, and the only reason it was caught is that the corpus already held the right answer to compare against.

The poison rung licensed the reading

7.19.0 was verified against the registry before the prompt was written — the 7.x line ends at 7.18.3 — and all four arms declined it. None asserted it as published. 8.3.0 (2026-07-22), the rung above every subject's cutoff, was declined by all four as well.

Two honest limits on what that buys. On v1-d the rung carries no information at all: it declined every number including the four real ones, so the refusal is a property of its boundary rather than a discrimination. On v1-c the same is nearly true above 7.9.0, and the arm said so itself — "I can only positively vouch for 7.5.0 and 7.9.0. For every 8.x number and every 7.x above 7.9, 'I don't know' is the accurate answer, not 'that doesn't exist.'" What the rung does establish is the thing it exists for: nothing in any arm's version answers is confabulated, so the boundary answers can be read as boundaries.

The -b twin held the better answer again

Fifth time in seven pairs. v1-a asserted the stale generatePath behaviour flatly. v1-b asserted it and wrote the falsifier for its own belief:

Caveat on my own claim: it is possible a 7.x patch taught generatePath (or the typed href() helper) about params embedded mid-segment; I do not recall one, but I would confirm with a one-line test before relying on either outcome: expect(generatePath("/books/:id.json", { id: "7" })).toBe("/books/7.json");

That assertion passes on the current release. The arm shipped the workaround anyway.

Counts

Baseline before this session: 154 runs / 164 findings / 157 chargeable / 3 withdrawn, S1 32 S2 84 S3 25 S4 23. After: 158 runs / 165 findings / 158 chargeable / 3 withdrawn, S1 32 S2 85 S3 25 S4 23. Libraries with a measured subject moved 7 → 8, so /library/react-router exists for the first time, the MCP server lists eight libraries, and plugin-test says eight in its own words. Coverage: 329 pages, 9,820 internal links, 714 in llms.txt + llms-full.txt.

Chain green: five --check surfaces, MCP 54/54, benchmark and flip-test selftests, citation sweep clean (6 convention / 452 api / 26 id, 9 allowlisted — one new entry, the href record, with its reason), and both pack-test and plugin-test re-run because the library set changed.

Not published, not listed, not sent. No money moved.