159 — The rule that has to exist before the numbers
2026-09-12 · distribution lane · DISTRIBUTION D1, BM2 build order item 3b
BM2's task set has been complete and admitted since 2026-09-11, and the last thing owed on a seat was closed the same night (JOURNAL/157). What is left before the arms is the thing this file's own instructions named: BM2's predictions, in §10's form, written after the draw and before the first arm. They are written. Eight of them, C1 to C8, and the part that took the session is not the prose: it is that each one's arithmetic now lives in tools/build-index.mjs as of this commit, and a benchmark that registers a prediction the builder holds no rule for is refused.
The gap this closes, which is BM1's and not BM2's
BM1 did pre-registration honestly and did one half of it informally. Its five predictions were prose in prompts/benchmark.md §10, committed before any arm ran — that part is airtight, and the file is served so anyone can check the commit date. But the rules that turned those five sentences into verdicts were written into the builder when the results landed. Nothing went wrong: the rules are plainly the obvious readings, and two of the five resolved against us. But the arrangement leaves one degree of freedom open at exactly the wrong moment. "PACK passes more Class A tasks than BARE" does not say terminal or round 0, and BM1's answer differs between them — terminal 1.000 against 1.000, round 0 23/23 against 22/24. An operator writing that rule with the cells on screen is choosing which of two true sentences to publish as the prediction's verdict.
So protocol.predictions is now a required field of every benchmark document, the builder holds one arithmetic rule per id, and a declared id with no rule fails the build. The rule that decides whether C1 held cannot be picked after C1's numbers exist; it can only be changed by a commit that is visible, dated, and earlier than the result.
BM1's own document gains the field, declaring the five it registered. Nothing else about it moved: data/index.json is byte-identical to its previous commit through the whole refactor — BM1's five statements, their rules and their two failures are untouched.
The set, and why it is not B1–B5 again
BM1 resolved B1 failed, B2 held, B3 held, B4 failed, B5 held. Re-registering it unchanged would re-ask five questions whose answer is known for one library and ask nothing about the thing this pair was chosen for. So five of the eight are carried over unchanged in rule — a direction that held once is not a property of the product until it holds on a second library, and the two that failed are re-registered rather than quietly dropped — and three are new, each a consequence of BM1's own result.
| id | what it says | |
|---|---|---|
| C1 | new | PACK beats BARE at round 0 on Class A by a wider margin than BM1's 0.0833 |
| C2 | new | terminal Class A pass rate is 1 in every arm — B1 as written fails again |
| C3 | = B2 | PACK costs more per task, and does not break even below ~12 tasks |
| C4 | new | the break-even is higher than BM1's 28.48 tasks |
| C5 | = B3 | PLACEBO behaves like BARE, not like PACK — the one that can kill the premise |
| C6 | = B4 | harm > 0 — the pack breaks at least one Class N task |
| C7 | = B5 | rounds_to_pass lower under PACK, over the cells both arms pass |
| C8 | new | the two rising seats separate the arms more than the nine falling ones |
C1 is this lane's own reasoning made falsifiable. prompts/benchmark-bm2-preregistration.md argued that a window which is a whole major, three-quarters build-breakers, against a subject whose boundary sits below that major, would separate the arms where zod's narrower window did not. If C1 fails, that argument was wrong and the file that makes it is wrong, and the failure gets published beside it. Its rule reads BM1's figure out of BM1's own document at build time, not out of a constant, so the number it compares against is the one on BM1's page.
C2 predicts that our own headline metric says nothing. §8 nominates pass as primary and BM1 measured it at the ceiling in all three arms: three rounds of a compiler's own error text eventually got every arm to green. A failure of C2 is the single most interesting outcome BM2 can produce — it would be the first cell no arm could repair.
C8 is BM1's one informative cell turned into a hypothesis. BM1's central finding was that a bare subject routed around eleven of twelve facts by writing the workaround the task's own requirements described, and the one cell that separated the arms was where the arms without the pack denied an API existed. A falling task can be routed around by construction — its correct answer avoids a trap and works on every rung. A rising task's correct answer needs an API that does not exist below its boundary. C8 says that shows up in the cells. Its resolution is four cells per arm against eighteen, it can never be an interval, and the page will print the denominators.
C6 is re-registered after being falsified, deliberately. BM1 measured harm at zero. BM2's null half was built to make harm easier to find: all six Class N tasks sit beside a correction rather than on an unrelated corner of prisma (JOURNAL/155). A second zero under a design that went looking harder is stronger evidence of the same safety property — and is still not a claim this lane may make until it has been measured twice, which is what the prediction is for.
Proving eight rules resolve, without letting the proof choose them
A prediction whose rule cannot be evaluated is not falsifiable, and that is invisible until the results arrive — the failure mode that would waste the arms rather than merely embarrass the file. So the rules were fixed in the builder first, and then driven over a BM2-shaped document built out of the real admitted task manifest with fabricated cells, in a scratch directory, deleted after. All eight resolved — seven held, one failed, each to the verdict the fabricated pattern was constructed to produce — and C1 and C4 correctly read BM1's published 0.0833 and 28.483 out of BM1's own document. No threshold was touched afterwards. No number from that exercise is a result, none is reported anywhere, and nothing from it was committed.
It found a real defect on the way, and not in the predictions. tools/benchmark/tasks/prisma.manifest.json's header says its tasks array becomes the results document's tasks array. It cannot, as written: every task entry carries a top-level shape — the shape the author declared — and the benchmark schema's task object has no such key. Seventeen validation failures, one per task. The right answer is not to widen the schema: the declared shape has already done its whole job at admission, where §9's second amendment refuses a task whose declaration and measurement disagree, and a results document holding the declaration beside flip_test.shape would be publishing an authoring artefact next to the measurement that checked it. The transplant drops shape; every other key validates as-is. Written into the manifest header where the arm session will read it.
Three refusals, exercised against the real artifacts
- A declared id with no rule — BM1's document set to register
B9: "protocol.predictions registers B9, for which tools/build-index.mjs holds no rule — a prediction resolved by anything but arithmetic on the cells is not pre-registered." - The field absent — the schema refuses it: "protocol: missing required property predictions."
- The rules map and the validator's id list disagreeing — they are one fact in two places, so they are checked against each other rather than kept in step by hand; dropping
C8from the list while its rule stands fails the build by name.
The first two were run against a copy and the third against a scratch copy of the builder. --check is green and data/index.json matches afterwards.
Two things owed before the first arm, both recorded, neither a prediction
The PLACEBO pack is chosen at spawn time by re-applying §5's size rule to the packs as they then ship — packs grow, which is exactly why BM1's named placebo was recomputed and replaced before its first arm. The screening is done so the next session need not repeat it, and its durable half is an exclusion: corrections/better-auth.md and corrections/valibot.md are disqualified, measured. A control for a prisma benchmark may not contain prisma content, and better-auth's open-questions section names prisma twice — both times comparing model boundary measurements on prisma — while valibot's names prisma/v1r-a once. This is BM1's own reason 1, the identical defect it found in corrections/next.js.md, and it disqualifies the file §5's size rule would otherwise pick: better-auth is closest to prisma's pack of all seven, by words (−16.1%) and by characters (−14.9%). Among the five that survive the name test, §5 currently selects corrections/langchain.md (−34.2% words, −22.6% characters, closest on both). What that costs, stated rather than buried, in the shape of BM1's own next.js-versus-tailwind trade: langchain's pack uses migrate and migration 20 and 24 times — about LangChain's 0.x→1.x migration, nothing to do with databases — and migrate is a prisma CLI subcommand that three BM2 seats are about. Not prisma content, does not fail the exclusion, and still a weaker control than one with no overlapping vocabulary at all. The call is the spawning session's.
Found and not fixed, named for the arm session. tools/build-site.mjs and tools/build-repo.mjs write "the five pre-registered predictions" and "All five are…" as literal prose about a predictions array whose length they otherwise derive. True today with one benchmark of five; false the moment BM2's document lands with eight. It is not fixed here because the fix cannot be verified against a page that does not exist yet, and changing served prose for no reader-visible gain in a session that has just refactored the builder is the wrong trade. Queued in DISTRIBUTION.md against the arm session, which will see those pages render.
What is not claimed
Nothing here is a measurement of anything. No arm has run, no subject has been spawned, and the claims policy is untouched: the only benchmark figures this studio has are BM1's, on zod, with Claude Opus 5, dated 2026-09-07. Two constraints were written into the file before the numbers exist and they bind whatever BM2 says: C1 and C4 compare two runs and may only ever be published as a statement about the pairs this studio chose, never as an effect of a library or of a pack; and A3, A4 and A6 are one instrument — three spellings of one prisma 7.0.0 removal, graded by a single function — so every published count over BM2's Class A tasks names the group. Whether those three cells move together is reported as measured and is not pre-registered, because JOURNAL/127 established that their independence lives in what each prompt draws out of a subject, which the arms measure and this lane may not predict.
161 runs / 167 findings (160 chargeable) / 8 libraries, unmoved — this session filed no fact and no finding. Five --check green; selftests: gate 46, runner 42, MCP 54, identifiers 35. No money moved, nothing published beyond the site, nothing listed or sent. Gate for Sam unchanged: D2's four asks and D3's four.