Does a correction pack change the code an AI writes? — measured on zod × Claude Opus 5, 2026-09-07

Pre-registered, three arms, every answer graded by running it. Two of the five predictions did not hold, and both are published here.

What was measured

One library (zod, graded against 4.5.4), one subject (Claude Opus 5), and 18 coding tasks: 12 whose correct answer depends on a fact this Index publishes about zod, and 6 null tasks on parts of the library that did not change — the only place the pack can be caught doing harm by applying a correction where none was needed. Each task was sent to a fresh subject 2 times in each of 3 arms: BARE, PACK, PLACEBO. PACK gets this library's correction pack in context; PLACEBO gets a correction pack of comparable length for a library the task never mentions, so a PACK effect can be told apart from the effect of having any long document in context.

Every answer is graded by executing an assertion against the installed release — no model judges another model's output. A failing answer is sent the toolchain's own error, in wording identical across arms, up to the pre-registered cap. The protocol, the metrics, the grading rules and the five predictions were committed before any arm ran; the task set was fixed by a recorded, recomputable draw before a single acceptance assertion was written.

Benchmarkzod--claude-opus-5--bm1--2026-09-07
SubjectClaude Opus 5 claude-opus-5
Cells106 graded of 108 planned (1 excluded, 1 abandoned)
Date2026-09-07

The result

ArmCells gradedPassedClass A passed at round 0 Correction roundsMean chars per cellRetry chars per cell
BARE3636/36 22/246 4,4961,799
PACK3535/35 23/230 53,7330
PLACEBO3535/35 20/237 74,21513,782

The metric the protocol nominated as primary did not separate the arms at all. Every arm passed every cell it was graded on, so the terminal pass rate on the fact-dependent tasks is BARE 100%, PACK 100%, PLACEBO 100% and says nothing: a retry loop that hands the model its own toolchain error lets all three converge. The separation that exists is at the first attempt, before any correction — BARE 22/24, PACK 23/23, PLACEBO 20/23.

13 correction rounds were spent in the whole run and 12 of them belong to one task, A6. 11 of the 12 fact-dependent tasks were answered correctly by every arm at the first attempt — a bare model routes around most of these facts, usually by writing the workaround the task's own requirements describe. That is a result rather than a defect in the draw: the admission gate was amended, before any arm ran, to refuse to require that a stale answer fail at the graded release, precisely so “the bare model routed around it” could be published.

The one task where the arms came apart

On A6 (fact LF17, boundary 4.1.13) the arms without the pack deny that the API exists, hold that position through correction rounds against the toolchain’s own output, and reach a passing answer only at the cap, by reverse-engineering the behaviour from what the failures leaked. With the pack the subject answers correctly at the first attempt, in both draws. Every reply quoted by that description is published below, unedited.

ArmRounds to pass
draw 1 / draw 2
Characters read
draw 1 / draw 2
Characters written back
draw 1 / draw 2
BARE 3 / 3 25,917 / 30,8756,451 / 8,662
PACK 0 / 0 53,229 / 53,229594 / 178
PLACEBO 3 / 3 264,199 / 257,37413,598 / 6,376

Characters, not tokens: the transport reports no provider token counts to the operator, this repository has no tokenizer and takes no dependency, and an estimated token count printed beside a measured one is a number nobody can check.

The subject’s replies for this task, every arm, every draw, every round:

What it costs

A PACK cell cost a mean of 53,733 characters against BARE’s 4,496, and BARE’s entire retry overhead is 1,799 characters per task. Billed the most generous way this instrument allows — the pack charged once per session, which is not how this run billed it — its 51,247 characters would need 28.5 fact-dependent tasks in one session to repay themselves against that overhead. The pre-registration predicted about 12. Billed per task, as this run billed it, never.

The most expensive arm is PLACEBO at 74,215 characters per cell: it pays for a long document on every round and still needs the rounds.

The five pre-registered predictions

Committed in the protocol before any arm ran, and resolved here by arithmetic on the cells rather than by the operator’s reading of them. The rule that decides each one is printed beside it.

#Pre-registered predictionRule that decides itMeasuredOutcome
B1PACK passes more Class A tasks than BARE. terminal Class A pass rate, PACK > BARE PACK 1, BARE 1 did not hold
B2PACK spends more per task than BARE, and does not break even below about 12 fact-dependent tasks per session. mean characters per cell, PACK > BARE, and pack context characters divided by BARE retry characters per cell is above 12 PACK 53,733, BARE 4,496, break_even_tasks 28.48 held
B3PLACEBO behaves like BARE, not like PACK, on Class A — so a PACK effect is the pack's content and not the bulk of any long text in context. round-0 Class A pass rate: PLACEBO is nearer BARE than PACK, ties counted as nearer BARE PLACEBO 0.87, BARE 0.92, PACK 1 held
B4The pack breaks at least one Class N task by over-applying a correction. harm above zero, where harm is a null task passing bare and failing with the pack harm 0, class_n_cells 12 did not hold
B5Rounds to pass are lower under PACK than BARE, over the tasks both arms eventually pass. mean rounds_to_pass over the cells both arms passed, PACK < BARE cells 35, BARE 0.17, PACK 0 held

Two did not hold, and both are against the studio rather than for it: the primary metric could not tell the arms apart, and the defect the protocol predicted in our own artifact did not appear in this sample — 12 null-task cells with the pack in context, none of them broken by it. A sample that small is quoted with its zero.

The per-task grid

The 12 fact-dependent tasks are not 12 independent trials. The draw is by fact and is blind to how facts cluster, and some drawn facts are answered together by one durable rule, so a subject that holds or acquires that rule moves several cells at once. The grid is published for that reason: the total would hide it.

TaskClassFactBoundary BARE
draw 1 / draw 2
PACK
draw 1 / draw 2
PLACEBO
draw 1 / draw 2
A1A LF3 4.4.0 pass / passpass / passpass / pass
A2A LF1 4.3.0 pass / passpass / passpass after 1 / pass
A3A LF10 4.4.0 pass / passpass / passpass / pass
A4A LF15 4.3.0 pass / passpass / passpass / pass
A5A LF9 4.4.0 pass / passpass / excludedpass / pass
A6A LF17 4.1.13 pass after 3 / pass after 3pass / passpass after 3 / pass after 3
A7A LF8 4.4.0 pass / passpass / passpass / pass
A8A LF7 4.2.0 pass / passpass / passpass / abandoned
A9A LF20 4.3.2 pass / passpass / passpass / pass
A10A LF18 4.4.0 pass / passpass / passpass / pass
A11A LF19 4.3.0 pass / passpass / passpass / pass
A12A LF12 4.4.0 pass / passpass / passpass / pass
N1N none — null task pass / passpass / passpass / pass
N2N none — null task pass / passpass / passpass / pass
N3N none — null task pass / passpass / passpass / pass
N4N none — null task pass / passpass / passpass / pass
N5N none — null task pass / passpass / passpass / pass
N6N none — null task pass / passpass / passpass / pass

The cells that are published and counted nowhere

2 of 108 planned cells enter no aggregate above. Each stays in the results file with its reply and its reason. A damaged or disqualified cell is repaired or abandoned, never re-drawn: a cell re-run after its result was seen is a cell chosen by its result.

What this does and does not say

What the pack bought, on this run: on the one fact-dependent task in 12 where the model denied that an API existed, the difference between answering at the first attempt and spending 3 and 3 rounds arguing with the toolchain. Nothing measurable on the other 11, at roughly 12 times the characters of a bare prompt on each. Narrow, expensive and real is the honest description.