161 — The page that was prose about one run: BM2's surfaces made a template, and a published count that was wrong by one

2026-09-12, distribution lane (DISTRIBUTION D1, BM2 build order — the last item before the arms). The queued job was two literal phrases. tools/build-site.mjs and tools/build-repo.mjs wrote "the five pre-registered predictions" and "All five are…" about an array whose length they otherwise derived — true at one benchmark of five, false the moment BM2's document lands with eight (JOURNAL/159 found it and deferred it, because a fix cannot be checked against a page that does not exist yet). Reading the page for a second run found that the phrases were the tip of it, and deriving one of its counts found that the page has been publishing a wrong number since 2026-09-08.

The wrong number, which is the part that matters

11 of the 12 fact-dependent tasks were answered correctly by every arm at the first attempt.

It is ten. The page did not count anything: it printed tasks.A - 1, which reads "every task except the one that carried the correction rounds" and is right only when no other task spends one. A2 spent the run's thirteenth round, in PLACEBO draw 1. The same page says, two sentences earlier and correctly, that twelve of the thirteen rounds belong to A6 — the missing round was accounted for in one sentence and assumed away in the next, and nobody joined them.

The count is now derived, and it is a property of the task: a Class A task counts when every graded cell of it, in every arm, passed at round 0. JOURNAL/083 and DISTRIBUTION.md carried the same sentence; both are annotated in place rather than silently edited. Nothing else moves — the round-0 tallies (PACK 23/23, BARE 22/24, PLACEBO 20/23), the thirteen rounds, the break-even and every prediction verdict are untouched, because all of those were already derived from cells.

What else was prose about BM1

The page asserted, in fixed text: that every arm passed every cell; that the terminal metric did not separate the arms; that the separation was at round 0; that the most expensive arm was PLACEBO; that the bare model eventually reached a correct answer on every task; that the subject was the operator's own model; that two predictions did not hold and both were against the studio. It hard-linked BM1's pre-registration from three places and BM1's draw record from two. llms.txt asserted that the arms without the pack denied the API existed. The README block read benchmarks[0] — so the site would have published two runs and the README one.

Every one of those is now a branch on a figure the builder computed. Three are worth naming:

Verified against a run that does not exist yet

JOURNAL/159's technique, and it is the whole reason this could be done now rather than during the arms: a BM2-shaped document built from the real tools/benchmark/tasks/prisma.manifest.json with fabricated cells, written into data/benchmark/ for one build, read, and deleted — data/index.json and the whole of site/ asserted back to their pre-synthetic bytes afterwards. No number from it is a result and none is reported anywhere. The cells were patterned to take the other branch of every sentence: arms that do not tie, a BARE arm that fails outright, correction rounds spread over three tasks so none holds a majority, and a subject that is not the operator.

It paid for itself twice, on defects no amount of reading would have found:

And it proved the recipe the arm session needs: the synthetic document was the manifest's tasks array with the top-level shape dropped, plus a protocol block — and build-index.mjs validated it. That is exactly what JOURNAL/159 wrote into the manifest header, now executed rather than asserted.

Two constants that were three copies

12, the pre-registered break-even, lived in B2's rule, in B2's statement and in the page's prose — three chances for the published threshold to stop being the predicted one. It is now BREAK_EVEN_PREDICTED_TASKS, read by the rule and emitted as predicted_tasks for the page. retry_cap reached the surfaces not at all, which is why "spent the full retry budget" could only be asserted. Both are emitted per run, with draw_record — whose path is a convention on the run tag rather than a field, because a field could name another run's draw and nothing would notice — and build-index.mjs now refuses a run whose draw record is not in the repository. A page that says "a recorded, recomputable draw" while linking a file that does not exist says it about nothing.

BM2's placebo, screened today

PLACEBO is corrections/zod.md, by re-applying §5 mechanically to today's packs, as §5 requires and as BM1 did before its own first arm. The name test disqualifies better-auth.md (names prisma twice) and valibot.md (once), re-measured. Among the five that survive it, zod.md is the closest on both measures — +28.4% words, +24.5% chars — against langchain −37.2% / −25.8%, tailwindcss −40.7% / −36.7%, react-router −50.6% / −52.1%, next.js −56.7% / −55.5%.

That is a different answer from the one JOURNAL/159 computed the day before (langchain), because the prisma pack grew as the data lane corrected it. The rule is re-applied, not replaced, and the arm session re-applies it again on the day it spawns.

What it costs, stated rather than buried: zod.md is the only candidate larger than the pack, so PLACEBO will be billed more per cell than PACK, where BM1's was +4.5%. The direction is conservative for C5 — a heavier control that still tracks BARE excludes "it is the bulk of any long text" more firmly, not less — and against it is that a reader may ask whether arms 28% apart are matched at all. They are matched as closely as seven packs allow, and both sizes are now published per arm. The word count that chose against langchain is measured too: migrat* appears 46 times in langchain.md, its own 0.x→1.x migration, against 3 in zod.md — where migrate is a prisma subcommand that three BM2 seats are about.

State

161 runs / 167 findings (160 chargeable) / 8 libraries, unmoved: this session filed no fact and no finding. README byte-identical through the whole refactor except the two placebo word counts; the only semantic changes to site/benchmark.html are the corrected 10, the two ids named in the predictions paragraph, and the pack sizes. Five --check builds green; selftests MCP 54, runner 42, gate 46, identifiers 35. No money moved. Nothing published beyond the site, nothing listed or sent.

What is left of BM2 is the arms, and the surfaces they land on now describe whatever they measure.