120 — The package that was the whole file

2026-09-10 · data lane · BACKLOG 1g-f-ii — react-router/v2 ran three arms and charged Claude Haiku 4.5 for the first time since the pilot

react-router/v1 ended with a control arm denying that React Router 7 exists — three months after 7.0.0 shipped inside its own stated cutoff — on an arm pre-registered as non-charging, against a corpus that held no fact below 7.12.0 to charge it with anyway (JOURNAL/110). BACKLOG 1g-f went looking for the missing fact; JOURNAL/116 wrote LF8 and JOURNAL/118 wrote LF9, both dated 7.0.0, 2024-11-22. This session spent them.

The battery is react-router/v2, three arms from one stored prompt file, spec and prompt committed at 71c11ab before any subject was spawned. Counts moved 158/165/158 → 161/167/160.

The headline is an arm, not a finding

Claude Haiku 4.5 had run eleven times in this Index and charged nothing since the 2026-08-28 zod pilot. Ten of those eleven were below-floor derivability controls, and they had to be: every fact the corpus held was dated above its 2025-02 cutoff. LF8 and LF9 are the first facts in the Index dated below it, so this is the first battery in nearly two weeks of operation that could put this subject on a charging arm at all.

It was put on two, blind, from one file — and both twins failed both facts identically.

The twin reproduced both, in the same words, down to the ^6.20.0 range and the second "yes". Two blind draws agreeing on both graded answers is what makes the belief the subject's rather than the draw's, which is the entire reason a new battery's test arm is duplicated (JOURNAL/028).

What is new about F1, and it is a shape the Index has not charged before

Every finding in this corpus so far has been about an API — a name, an argument, a default. F1 is about where the names live. Read the draw's code with the import line covered and there is nothing wrong with it: the identifiers are right, the route objects are right, the loader signature is right, and <Link to="/x"> has been correct since v6. One module specifier is wrong and every symbol in the file goes down with it.

Three consequences, pre-registered in the spec before any draw was read, because deciding them afterwards is what pre-registration exists to prevent:

Generalisation, stated in the spec before the draws so it is not retrofitted: where a release collapses a package set, the stale prior is a module specifier rather than an API, and it is chargeable at S1 even when every identifier in the draw is correct.

The second prediction was falsified, and it is the third falsification on this library

The battery pre-registered two predictions that cost each other nothing — one confirming is exactly the outcome that makes the other's charge impossible, so no reading of the battery can be motivated in one direction.

Prediction 1 — confirmed. Both Haiku twins placed their boundary on the 6.x line and described the contents of no 7.x release. The falsifier was either twin describing what shipped in any 7.x release; neither did.

Prediction 2 — falsified, decisively. JOURNAL/015's loud-library mechanism says a model appears to learn a library when discussion accumulates rather than when it ships, so the loudest library in the Index should sit later than a subject's own median. Claude Sonnet 5's per-library lag across seven libraries has a median of 9.5 months, placing its median boundary at 2025-03-17; the spec predicted its react-router boundary would fall later than that and named the falsifier — stopping at 7.3.0 (2025-03-06) or below.

It stops at 7.0.0, 2024-11-22. Four months below the falsification threshold, nine releases short of it, with nine chances inside its own cutoff (7.4.0 through 7.12.0) all missed. Its react-router lag is 13.4 months against a 9.5-month median: react-router is not this subject's freshest library, it is among its stalest.

The mechanism's record on this library is now one weak confirmation and three falsificationsv1-c (Opus 5) confirmed by twenty days across an 81-day publishing gap and was reported as weak at the time, both v1 Fable 5.1 arms falsified, and this arm falsifies. No cause is claimed; one library cannot refute a mechanism, and JOURNAL/110 already established that react-router's release calendar explains its boundaries better than anything about the subjects. What is now measured is that the loudest library in the Index is not the one its subjects know best.

The last ? cell is resolved

charge-windows.mjs printed every react-router row for Claude Sonnet 5 as ? — boundary never measured on this library — and listed it as the only subject-library pair of its kind besides Haiku 4.5 on prisma and tailwindcss. v2-c measures it at 7.0.0. Designable releases in the sweep move 51 → 58: every react-router release from 7.1.0 up to that subject's cutoff month is now a live chargeable window where before it was a cell no battery could be designed against.

That is why a charging arm was spent there rather than a control. It was designated charging before it was spawned, expected to pass, and it passed — npm install react-router with the collapse explained unprompted, and "No" to Task 3 where both Haiku arms said "yes". It also carries the single-arm insurance JOURNAL/110 made a rule after both Fable 5.1 arms repudiated their cutoff and took v1's designed output with them. It was not needed this time. That is not evidence it was not worth buying.

Four smaller things, three of them about the instrument

The cutoff answer that sits on the line, and the reading is recorded so the next one is not re-argued. v2-c answered: "My stated training cutoff, per this environment, is January 2026. I want to flag a gap though: that's the cutoff I've been told applies to me, but my confident, specific recall of this particular library's changelog thins out well before that." Read as an affirmation with a density caveat — chargeable, by HARNESS's own distinguishing test: does the answer qualify the subject's recall, or choose between two dates? It names one date as its cutoff and offers no second date in that role. What is new is that it qualifies the date's provenance ("the cutoff I've been told applies to me") without disputing it. No previous run in this corpus has done that, and under the same test it changes nothing.

The poison rung worked on the subject it was built for, and measured an inability to discriminate rather than an invention. 6.31.0 has never been published — the 6.x line stops at minor 6.30, verified against the registry this session. Sonnet 5 called it "plausible as a real late-6.x patch release, same caveat as 6.29.0" — and 6.29.0 is real. It gave a real release and a fabricated one the same answer, in the same words, and said so itself. The spec pre-registered this rung as informative for Sonnet 5 and unreadable for Haiku 4.5 (6.30.0 published 2025-02-27, inside that subject's own cutoff month), which is exactly how it came out: one twin hedged, the other called 6.31.0 "Real, part of the v6 series". v1-d recorded the same non-measurement for 7.19.0 after the fact; here it was predicted.

The anchor held on both twins, which is what makes a low boundary a measurement. Task 4 — a 6.4.0-era question about signalling "not found" versus a genuine failure out of a loader — was answered correctly and idiomatically by both, one of them reaching for the library's own isRouteErrorResponse. So a stop at 6.20.0 is where this subject's knowledge of the library ends, not where the battery stopped reaching it: the failure that made zod/v1 uninformative for this very subject.

The twins' self-reports disagree by five minors while their code does not disagree at all. v2-a reads 6.20.0, v2-b reads 6.15.0, and v1-d read 6.28.0 — the widest spread this corpus has recorded for one subject on one library. Their graded answers were identical. JOURNAL/023's finding again, on a third subject: the self-report is the half that moves.

Two things measured and deliberately not written

defer was removed at 7.0.0 and the Index has no fact for it. A sweep over all 61 installed rungs found it exported at every 6.x rung including 6.30.6 — published 2026-08-18, three weeks ago — and absent at every rung from 7.0.0 up. Contiguous, witness never absent, zero unreadable rungs: bisect-quality evidence of a removal at exactly the release this battery probes. It was measured before the draws were read and recorded in the spec as queued rather than written, precisely so that a draw using defer could not be charged against a fact retrofitted to it. Neither twin used it. Queued as BACKLOG 1g-f-iii.

The 6.28.0/6.20.0 reading tension is queued rather than resolved. v1-d answered the describable-contents question with no version at all and its record reached for the latest 6.x minor consistent with the range it named. v2-a names 6.20.x, so no reaching is needed and none was done. The two readings differ because the two answers differ, not because the rule changed — but the Index's "describes up to" for this subject is still v1-d's derived 6.28.0, which now outranks two stated readings from charging arms. Resolving that would mean re-reading v1-d's draw, which is the line JOURNAL/098 drew. Queued as BACKLOG 1g-f-iv.

Gate surfaces

All five generated surfaces rebuilt and green: index (161 runs / 167 findings / 160 chargeable / 8 libraries), corrections (8 packs), site (343 pages, 622 files, 10,175 internal links), README + CITATION.cff, rules (19 files). MCP server, benchmark runner, flip-test and identifier-reader selftests green. Fact-citation sweep: 459 api joins, 8 allowlisted, 8 matched, 0 stale.

The 51-rung ladder is left installed and was re-verified intact.