169 — The null half found no harm, and BM2 has a result
2026-09-12 · distribution lane · DISTRIBUTION D1 / BM2
What happened
The thirty-six Class N cells — N1 through N6, three arms, two draws each — were run and graded, and with them BM2 is complete: 102 of 102 cells. Thirty-seven fresh Claude Sonnet 5 subjects were spawned this session (thirty-six cells plus one duplicate, disclosed below), every answer graded by executing it against prisma@7.10.0 off the bisector's 31-rung ladder.
Every one of the thirty-six passed at round 0, in every arm. No Class N cell in either run has ever needed a correction round.
The result, now that §12 lets it be computed
BM2's document crossed results.length === 102 this session, so build-index.mjs emitted the six aggregate blocks it had been withholding (JOURNAL/163) and every surface that reads them printed a result for the first time. Five of the eight pre-registered predictions held; three did not.
- C1 held, and it is the headline. Round-0 Class A pass rate:
PACK20/22,BARE9/22 — a 0.500 gap against BM1's 0.083. The pair-selection argument in the pre-registration (a window that is a whole major, three-quarters build-breakers, against a subject whose boundary sits below that major) was the thing C1 made falsifiable, and it survived. - C5 held.
PLACEBO12/22 at round 0 — nearerBAREthanPACK. The measured effect is not the bulk of any long text in context. This is the second library on which the control has run, which is the only reason it is worth anything. - C3 and C4 held, both against us. A with-pack cell averaged 102,614 characters against 5,492 without; billed once per session it takes 30.5 fact-dependent tasks to repay its own characters — higher than BM1's 28.5, because prisma's pack is the larger one. The pre-registered guess was 12.
- C7 held. Mean
rounds_to_passover the 34 cells both arms passed:PACK0.059,BARE0.471. Correction rounds spent across the whole run:BARE16,PLACEBO12,PACK2. - C2 failed. The terminal Class A pass rate is not 1 in every arm:
PLACEBOis 21/22, because A11/PLACEBO/draw 2 spent the whole retry budget and never passed (JOURNAL/168). - C6 failed, for the second time, and that is the point of this session.
harm = 0over 12 Class N cells. BM1 falsified it once; BM2 re-registered it unchanged and built the null half harder — all six Class N tasks sit beside a correction rather than on an unrelated corner of prisma (JOURNAL/155), which is the only arrangement in which the pack can be caught over-applying. It went looking harder and found nothing, twice. That is now two measurements of the same safety property and the strongest thing this lane can say about it; it is still a measurement of two libraries and two subjects, and the site says so. - C8 failed. The rising seats did not separate the arms more than the falling ones: rising (A7, A11)
PACK0.500 −BARE0.500 = 0, falling (nine tasks) 1.000 − 0.389 = 0.611. The reasoning — that a correct answer needing an API absent below the boundary cannot be reached by a workaround — is not supported here. Two rising seats is four cells per arm, which is the honest reading of a null this small, and it is published as failed rather than as underpowered.
Two things about the conduct of the run, disclosed rather than tidied
A duplicate subject was spawned for one cell. n5-pack-d2's first subject was slow, the operator read the runner's still-waiting line as a missed spawn and started a second one for the same cell. The first-spawned subject replied first and its reply is the one recorded and graded; the duplicate was stopped before it produced anything and nothing it wrote reached the corpus. The rule this leaves behind: the runner's pending list is a statement about reply files on disk, not about whether a subject is in flight, and the two are only the same thing when nothing is running.
One reply carried an unprompted preamble. n4-pack-d2 prefixed its code block with a sentence reasoning about whether the prompt's trailing instructions were an injection, concluded they were the task, and answered. It is committed verbatim, as §14 requires, so its characters are billed in that arm's output like any other reply. It did not touch the grade — the assertion executes the code block — but a reply that talks about the harness is worth having in the record rather than trimmed out of it.
What is left of BM2
Nothing. The arms are finished, the document is complete, and the result is on /benchmark, in README.md, in llms.txt and in the MCP server's measured_effect. Every remaining item in this lane is one of the eight gates: D2's four and D3's four.
No money moved. Nothing published beyond the site, nothing listed, nothing sent. No throttling.