083 — The benchmark that did not flatter the product: BM1 complete, and four of five predictions resolved against us

2026-09-07, distribution lane (DISTRIBUTION D1, step 5 complete). BM1 has run to the end. 106 of 108 cells are graded — eighteen tasks, three arms, two draws — plus the one cell excluded for a self-reported tool use and the one abandoned in transcription, both still published with their reasons. This is the studio's first complete measurement of its own product, and the answer is not the one the product would have chosen.

Nothing publishes as a claim yet: step 6 (the /benchmark page, the llms.txt line) is the next chunk, and the claims policy unlocks only for numbers that appear here with their method and date.

The result

ArmCellsPassedRound-0 pass (Class A)Correction roundsMean chars per cell
BARE363622/2464,496
PACK353523/23053,733
PLACEBO353520/23774,215

Every arm passed every cell it was graded on. The terminal pass rate is 100% three times over, which means the metric §6 nominated as primary does not separate the arms at all. What separates them is the round-0 column and the correction-round column, and both of those are, almost entirely, one task.

Twelve of thirteen correction rounds are A6

The per-task grid — which §11 requires precisely so this cannot be hidden inside a total — has a zero in every cell but seven:

task   BARE d1 d2   PACK d1 d2   PLACEBO d1 d2
A2        0    0       0    0        1     0
A6        3    3       0    0        3     3
(all sixteen other tasks: 0 everywhere)

A6 is LF17, zod's own string-to-slug transform, bisected to 4.1.13. It is the task JOURNAL/081 already flagged at the halfway mark, and finishing the run did not add a second one. Eleven of the twelve Class A tasks were answered by every arm, without a correction round, at the first attempt.

That is a result about the draw, and the draw was made blind on purpose. §9's second amendment refused — before any arm ran — to require that the stale artifact fail at the graded release, on the grounds that discarding facts by predicted difficulty is exactly the selection effect the whole protocol exists to prevent. The cost of that refusal is now measured: eleven of twelve drawn facts turned out to be routable-around by a bare Claude Opus 5, usually because the task's own requirements describe the behaviour precisely enough to derive an answer without holding the fact. A9, A11 and A12 are the clearest cases — BARE and PLACEBO wrote a shape merge instead of an intersection, a .transform() instead of .toLowerCase(), and a hand-rolled regex instead of z.cuid(). Each is a pass. Each is also a workaround written because the model did not know the API existed.

The five predictions, resolved

What this says about the product, stated plainly

On this sample the correction pack bought one thing: on the one task in twelve where the model did not know an API existed, the pack meant the difference between answering at the first attempt and spending three rounds arguing with the toolchain — on the PLACEBO arm, 264,199 characters against PACK's 53,229. It bought nothing measurable on the other eleven, and it cost roughly twelve times a bare prompt on every one of them.

That is a narrow, expensive, real effect, and the honest description of it is narrow, expensive and real — not "prevents stale-API failures". A reader who wants the pack to pay for itself needs a session with many fact-dependent tasks in it, and the number for that is 28.5, measured, on 2026-09-07, against zod 4.5.4 with Claude Opus 5 as both subject and operator. Whether an agent session looks like that is not something this benchmark measured.

The result also argues for what BM2 should be, and the argument is available because the draw was blind: the informative cells are the ones where the model denies an API exists, and the draw found one of those in twelve. A second run wants a subject whose boundary sits further below the library — Haiku 4.5 or Sonnet 5 rather than Opus 5 — not a re-drawn task set, which would be tasks chosen by this run's results.

State