Pre-registered, three arms, every answer graded by running it. Two of the five predictions did not hold, and both are published here.
One library (zod, graded against
4.5.4), one subject (Claude Opus 5), and
18 coding tasks: 12 whose correct answer depends on a fact this
Index publishes about zod, and 6 null tasks on parts of the
library that did not change — the only place the pack can be caught doing harm by applying a
correction where none was needed. Each task was sent to a fresh subject 2 times in
each of 3 arms: BARE, PACK, PLACEBO.
PACK gets this library's correction pack in context; PLACEBO gets a
correction pack of comparable length for a library the task never mentions, so a PACK
effect can be told apart from the effect of having any long document in context.
Every answer is graded by executing an assertion against the installed release — no model judges another model's output. A failing answer is sent the toolchain's own error, in wording identical across arms, up to the pre-registered cap. The protocol, the metrics, the grading rules and the five predictions were committed before any arm ran; the task set was fixed by a recorded, recomputable draw before a single acceptance assertion was written.
| Benchmark | zod--claude-opus-5--bm1--2026-09-07 |
|---|---|
| Subject | Claude Opus 5 claude-opus-5 |
| Cells | 106 graded of 108 planned (1 excluded, 1 abandoned) |
| Date | 2026-09-07 |
| Arm | Cells graded | Passed | Class A passed at round 0 | Correction rounds | Mean chars per cell | Retry chars per cell |
|---|---|---|---|---|---|---|
BARE | 36 | 36/36 | 22/24 | 6 | 4,496 | 1,799 |
PACK | 35 | 35/35 | 23/23 | 0 | 53,733 | 0 |
PLACEBO | 35 | 35/35 | 20/23 | 7 | 74,215 | 13,782 |
The metric the protocol nominated as primary did not separate the arms at all. Every arm passed every cell it was graded on, so the terminal pass rate on the fact-dependent tasks is BARE 100%, PACK 100%, PLACEBO 100% and says nothing: a retry loop that hands the model its own toolchain error lets all three converge. The separation that exists is at the first attempt, before any correction — BARE 22/24, PACK 23/23, PLACEBO 20/23.
13 correction rounds were spent in the whole run and 12 of them belong
to one task, A6. 11 of the 12
fact-dependent tasks were answered correctly by every arm at the first attempt — a bare model
routes around most of these facts, usually by writing the workaround the task's own requirements
describe. That is a result rather than a defect in the draw: the admission gate was amended, before
any arm ran, to refuse to require that a stale answer fail at the graded release, precisely so
“the bare model routed around it” could be published.
On A6 (fact LF17, boundary
4.1.13) the arms without the pack deny that the API exists, hold that
position through correction rounds against the toolchain’s own output, and reach a passing
answer only at the cap, by reverse-engineering the behaviour from what the failures leaked. With
the pack the subject answers correctly at the first attempt, in both draws. Every reply quoted by
that description is published below, unedited.
| Arm | Rounds to pass draw 1 / draw 2 |
Characters read draw 1 / draw 2 |
Characters written back draw 1 / draw 2 |
|---|---|---|---|
BARE |
3 / 3 | 25,917 / 30,875 | 6,451 / 8,662 |
PACK |
0 / 0 | 53,229 / 53,229 | 594 / 178 |
PLACEBO |
3 / 3 | 264,199 / 257,374 | 13,598 / 6,376 |
Characters, not tokens: the transport reports no provider token counts to the operator, this repository has no tokenizer and takes no dependency, and an estimated token count printed beside a measured one is a number nobody can check.
The subject’s replies for this task, every arm, every draw, every round:
A PACK cell cost a mean of 53,733 characters against
BARE’s 4,496, and BARE’s entire retry overhead is 1,799
characters per task. Billed the most generous way this instrument allows — the pack charged
once per session, which is not how this run billed it — its
51,247 characters would need 28.5 fact-dependent
tasks in one session to repay themselves against that overhead. The pre-registration
predicted about 12. Billed per task, as this run billed it, never.
The most expensive arm is PLACEBO at 74,215 characters
per cell: it pays for a long document on every round and still needs the rounds.
Committed in the protocol before any arm ran, and resolved here by arithmetic on the cells rather than by the operator’s reading of them. The rule that decides each one is printed beside it.
| # | Pre-registered prediction | Rule that decides it | Measured | Outcome |
|---|---|---|---|---|
B1 | PACK passes more Class A tasks than BARE. | terminal Class A pass rate, PACK > BARE | PACK 1, BARE 1 | did not hold |
B2 | PACK spends more per task than BARE, and does not break even below about 12 fact-dependent tasks per session. | mean characters per cell, PACK > BARE, and pack context characters divided by BARE retry characters per cell is above 12 | PACK 53,733, BARE 4,496, break_even_tasks 28.48 | held |
B3 | PLACEBO behaves like BARE, not like PACK, on Class A — so a PACK effect is the pack's content and not the bulk of any long text in context. | round-0 Class A pass rate: PLACEBO is nearer BARE than PACK, ties counted as nearer BARE | PLACEBO 0.87, BARE 0.92, PACK 1 | held |
B4 | The pack breaks at least one Class N task by over-applying a correction. | harm above zero, where harm is a null task passing bare and failing with the pack | harm 0, class_n_cells 12 | did not hold |
B5 | Rounds to pass are lower under PACK than BARE, over the tasks both arms eventually pass. | mean rounds_to_pass over the cells both arms passed, PACK < BARE | cells 35, BARE 0.17, PACK 0 | held |
Two did not hold, and both are against the studio rather than for it: the primary metric could not tell the arms apart, and the defect the protocol predicted in our own artifact did not appear in this sample — 12 null-task cells with the pack in context, none of them broken by it. A sample that small is quoted with its zero.
The 12 fact-dependent tasks are not 12 independent trials. The draw is by fact and is blind to how facts cluster, and some drawn facts are answered together by one durable rule, so a subject that holds or acquires that rule moves several cells at once. The grid is published for that reason: the total would hide it.
| Task | Class | Fact | Boundary | BARE draw 1 / draw 2 | PACK draw 1 / draw 2 | PLACEBO draw 1 / draw 2 |
|---|---|---|---|---|---|---|
A1 | A | LF3 | 4.4.0 |
pass / pass | pass / pass | pass / pass |
A2 | A | LF1 | 4.3.0 |
pass / pass | pass / pass | pass after 1 / pass |
A3 | A | LF10 | 4.4.0 |
pass / pass | pass / pass | pass / pass |
A4 | A | LF15 | 4.3.0 |
pass / pass | pass / pass | pass / pass |
A5 | A | LF9 | 4.4.0 |
pass / pass | pass / excluded | pass / pass |
A6 | A | LF17 | 4.1.13 |
pass after 3 / pass after 3 | pass / pass | pass after 3 / pass after 3 |
A7 | A | LF8 | 4.4.0 |
pass / pass | pass / pass | pass / pass |
A8 | A | LF7 | 4.2.0 |
pass / pass | pass / pass | pass / abandoned |
A9 | A | LF20 | 4.3.2 |
pass / pass | pass / pass | pass / pass |
A10 | A | LF18 | 4.4.0 |
pass / pass | pass / pass | pass / pass |
A11 | A | LF19 | 4.3.0 |
pass / pass | pass / pass | pass / pass |
A12 | A | LF12 | 4.4.0 |
pass / pass | pass / pass | pass / pass |
N1 | N | none — null task | — | pass / pass | pass / pass | pass / pass |
N2 | N | none — null task | — | pass / pass | pass / pass | pass / pass |
N3 | N | none — null task | — | pass / pass | pass / pass | pass / pass |
N4 | N | none — null task | — | pass / pass | pass / pass | pass / pass |
N5 | N | none — null task | — | pass / pass | pass / pass | pass / pass |
N6 | N | none — null task | — | pass / pass | pass / pass | pass / pass |
2 of 108 planned cells enter no aggregate above. Each stays in the results file with its reply and its reason. A damaged or disqualified cell is repaired or abandoned, never re-drawn: a cell re-run after its result was seen is a cell chosen by its result.
A5 / PACK / draw 2 —
excluded. Excluded from every aggregate under prompts/benchmark.md §14: the subject self-reported TOOLS_USED: Bash, so it did not answer under the no-tools condition every other cell answered under. The cell stays published with its reply and its verdict (pass, round 0). The operator cannot see a subagent tool call, so this exclusion rests on the subject having reported it.A8 / PLACEBO / draw 2 —
abandoned. Abandoned, not graded, counted in no aggregate. The subject replied and the reply reached the operator, but saving it introduced a U+25A0 inside the object literal - operator transcription damage, not the subject's code - and node read it as SyntaxError: Invalid or unexpected token. The operator cannot reconstruct the bytes the subject sent, so the cell is not evidence either way and is not re-run: a cell re-drawn after its result was seen is a cell chosen by its result. The damaged reply stays on disk as prompts/sent/benchmark/bm1/a8-placebo-d2.round0.reply.txt. tools/benchmark/run.mjs now refuses to grade a reply carrying that class of character (prompts/benchmark.md section 14).What the pack bought, on this run: on the one fact-dependent task in 12 where the model denied that an API existed, the difference between answering at the first attempt and spending 3 and 3 rounds arguing with the toolchain. Nothing measurable on the other 11, at roughly 12 times the characters of a bare prompt on each. Narrow, expensive and real is the honest description.