# The benchmark — does the correction pack change the outcome?

**Pre-registration. Written and committed 2026-09-07, before any arm has run.** DISTRIBUTION D1;
Decision 009. This file is the protocol. The task set (§7) and the sent prompts go in
`prompts/sent/benchmark-<lib>-<n>.txt` and are committed before spawning, under the standing rule
(HARNESS § *Write and commit the battery spec BEFORE spawning*).

## 1. Why this exists, and what it gates

The claims policy (DISTRIBUTION § *Claims policy*) forbids every benefit claim — "saves tokens",
"prevents failures", "worth the money" — anywhere public until this instrument has a number. Nothing
unlocks before it. D8, the paid tier, is downstream of it. So this is the highest-value thing the
distribution lane can build, and it is also the thing most able to embarrass us: it is designed so
that a result showing the pack does **nothing** is publishable in exactly the same shape as a result
showing it works.

The studio has 68 runs measuring *what models believe*. Not one measures *whether the correction
changes what a model produces*. That gap is what a buyer, and an AI recommending a purchase, would
ask about first.

## 2. What this instrument is, and what it is not

**It is:** a measurement of whether placing `corrections/<lib>.md` in a model's context changes the
code that model writes, graded by executing that code against the installed current release.

**It is not** a measurement of the model's knowledge — that is what the batteries do. A benchmark arm
elicits code to be *run*, not beliefs to be *read*. Its runs are not batteries, carry no findings, and
never enter `runs_eliciting_code` or the correction packs. They live in `data/benchmark/`, under their
own schema, and are reported only on `/benchmark`.

**It is not fair to the pack, deliberately.** The pack states the answer to any task derived from one
of its facts. On those tasks the corrected arm is open-book. That is the product's actual mechanism —
the pack is *meant* to be in context — but it means a win on Class A tasks is weak evidence and must
never be reported as though the model got smarter. The load-bearing measurements are the ones that can
go against us: the placebo arm (§5), the null tasks (§7), and the token ledger (§8).

## 3. The question, as a measurable

> For a task whose correct answer depends on a library fact that a model does not hold, does adding
> that library's correction pack to context change (a) whether the produced code passes an executable
> acceptance check, (b) how many correction rounds it takes to pass, and (c) the total tokens spent
> reaching a pass — and does the pack cost more than it saves?

Every one of (a), (b), (c) can come back null or negative. All three publish identically.

## 4. Scope of v1

- **Library:** `zod`. Chosen because it is the only library whose facts have been **bisected** —
  every `introduced_in` confirmed by executing the fact's own statement against 15 installed releases
  (JOURNAL/067, `tools/audit/probes/zod.mjs`). Eight of its 23 dated facts were wrong before that
  sweep. A benchmark task whose "correct answer" rests on an unverified release note measures the
  release note. No library enters this benchmark before its facts are bisected.
- **Subject, BM1:** Claude Opus 5. Largest eligible pool (16 facts, §7), 58 runs of prior
  characterisation, boundary `zod 4.1.0` stable across the corpus.
- **Later:** Claude Sonnet 5 and Claude Fable 5.1 on zod, then a second bisected library. Extending
  the subject list does not change this protocol; extending the metric list does, and requires a new
  pre-registration section dated before those arms run.

## 5. Arms

Three arms, each drawn **in duplicate** (HARNESS § *Every new battery runs its test arm in duplicate*
— the duplication protects a claim, and this instrument exists to produce one), blind and concurrent,
from one stored prompt file per arm.

| Arm | Context prepended to the task | Purpose |
|---|---|---|
| `BARE` | nothing | the counterfactual |
| `PACK` | `corrections/zod.md` (5,384 words) | the product |
| `PLACEBO` | `corrections/next.js.md` (5,560 words) | the active control |

**The placebo is the whole design.** `PACK` differs from `BARE` in two ways at once: it adds the right
content, *and* it adds ~8k tokens of dense, dated, house-style library prose that may sharpen attention
on library-version questions regardless of content. `PLACEBO` holds the second constant and removes the
first. It is a real correction pack for a library the task never mentions — no fabricated text is
generated for this experiment, ever. Word-count mismatch is 3.3% (next.js 5,560 vs zod 5,384),
the closest match among the seven packs; the exact token counts of all three contexts are recorded
per run rather than assumed.

Subjects cannot be blinded — a model can see a pack in its own context. They do not need to be: the
grader is `node`, not a reader (§9).

**Amendment, 2026-09-07, written before any arm ran: the placebo is `corrections/tailwindcss.md`,
not `corrections/next.js.md`. Two reasons, and the first one alone is disqualifying.**

1. **`corrections/next.js.md` names zod three times.** Once inside a technical explanation
   (`normalizeNextConfigZodErrors`, the mechanism by which an unrecognised `images` key exits the
   build) and twice in its open-questions section, where the entries compare **model boundaries on
   zod** to model boundaries on next.js. A control whose whole job is to hold context bulk constant
   while removing zod content cannot contain zod content — least of all a sentence about a subject's
   zod boundary. This was not true when §5 was written, or it was not checked; it is true of the
   file as it stands today, and it was found by grepping the candidate before the arm was built.
   `corrections/tailwindcss.md` contains no occurrence of `zod`, `schema`, `refine`, `coerce`,
   `discriminatedUnion` or `pick(` — one mention of `valibot`, in a cross-reference to a defect
   class, and nothing about schema validation at all.
2. **The size rule §5 states now selects a different file.** §5 chose next.js as "the closest match
   among the seven packs" at 5,560 words against zod's 5,384. The zod pack has since grown to
   **7,396 words / 51,247 chars** — the data lane's corrections, which is what the product is
   supposed to do. Re-applying §5's own rule, mechanically, over the six non-zod packs:
   tailwindcss 7,732 words (**+4.5%**, 57,591 chars), langchain 8,082 (+9.3%), prisma 8,889
   (+20.2%), better-auth 9,687 (+31.0%), next.js 5,560 (**−24.8%**), valibot 4,310 (−41.7%).
   Tailwind is closest by words and by characters. The rule is re-applied, not replaced.

What this costs, stated rather than buried: next.js is *adjacent* to a zod task — server-side
TypeScript — and tailwind is not, so a placebo drawn from CSS is a weaker test of a priming effect
that depends on domain adjacency than one drawn from the same neighbourhood. That is a real
argument for next.js and it loses to reason 1, because a control containing the tested library's
name and a claim about its boundaries is not a control. The trade is recorded here so a reader of
B3 can weigh it, and the arm's exact bytes are pinned by `context_sha256` in the results file.

The zod pack's growth is itself recorded: the arms are graded against the pack **as it ships today**
(sha256 `69fdf7ae…`), not the 5,384-word version §5 was written against. A benchmark of a stale copy
of our own product would measure nothing anyone can buy.

## 6. Task classes

**Class A — charging.** Correct answer depends on exactly one eligible fact (§7). Prediction: `BARE`
fails, `PACK` passes.

**Class N — null.** An idiomatic task on the same library whose correct answer depends on **no** fact
in the pack: it uses only APIs whose behaviour is identical across every rung of the zod ladder
(`4.0.0` … `4.5.4`), which is verified by running the task's own acceptance check against all fifteen
installed releases before the task is admitted. Prediction: every arm passes.

Class N is not filler. It is the only place the pack can be caught doing **harm** — refusing a fine
API, pinning a version nobody asked about, rewriting working code toward a correction that does not
apply. A pack that fixes Class A and breaks Class N is a worse product than no pack, and nothing in
the Class A numbers would show it.

Ratio: **2 Class A : 1 Class N**, fixed before task authoring, so the null tasks cannot be quietly
dropped if they misbehave.

## 7. Eligibility, and the pool as it stood on 2026-09-07

A fact may source a Class A task only if all four hold:

1. **Bisected.** `introduced_in` confirmed by an executable probe, not by a release note.
2. **Above the subject's measured boundary** for this library (`charge-windows.mjs`), so the subject
   is not expected to hold it.
3. **Below the subject's stated cutoff**, and not in the cutoff month — the standing fairness rule
   (HARNESS § *The subject's stated cutoff is a draw too*). A parked release charges nothing.
4. **Executably decidable** against the installed current release (§9).

Computed for zod on 2026-09-07, before any task was written:

| Subject | Boundary | Eligible facts | Throws | Assert-only |
|---|---|---|---|---|
| Claude Sonnet 5 | > 4.0.0 | 11 | LF1, LF2 | 9 |
| **Claude Opus 5** | **> 4.1.0** | **16** | **LF1, LF3** | **14** |
| Claude Fable 5.1 | > 4.2.0 | 12 | LF1, LF3 | 10 |

Eligible for all three subjects: **LF1, LF4, LF15, LF19, LF20.**

**The finding that changed this design.** D1 as written in DISTRIBUTION.md names the primary metric
"stale-API failures (code executed against the installed current version)". Read strictly — code that
*fails to execute* — the pool is **two facts per subject**. Everything else is `added`,
`behavior-changed` or `stricter`: the stale belief there produces code that runs and is wrong, or code
that works and is worse. Two facts cannot support a benchmark, and a benchmark built only on the two
would measure the rarest failure mode and overstate the rest by silence. So the primary metric is
widened, here and before any arm ran, from *throws* to *fails its executable acceptance assertion*
(§8). D1's wording is amended to match. Recording this now is the difference between a design decision
and a post-hoc one.

**Selection rule, to keep the corpus from choosing the tasks.** Class A facts are drawn from the
eligible pool **stratified by severity** (S1, S2, S3 in the pool's own proportion), *without* reference
to which facts prior batteries already charged against this subject. Building the benchmark out of
known failures would inflate `BARE`'s failure rate by construction and produce a number we could not
honestly quote. Both cuts are reported: the whole sample, and the previously-charged subset, separately.

## 8. Metrics

Per arm, per task, recorded in `data/benchmark/<lib>-<subject>-<n>.json`:

- **`pass`** *(primary)* — the produced code satisfies the task's executable acceptance assertion when
  run against the installed current release. Boolean, machine-set.
- **`rounds_to_pass`** — correction rounds needed, `0` if it passes first try, `null` if never
  (§9 cap). The loop feeds back only what a real toolchain emits: the `tsc` diagnostic or the thrown
  error, verbatim. No hints, no restatement of the task, identical wording across arms.
- **`tokens`** — input and output separately, per round, summed. **The pack's own tokens count against
  it in every round it is resent**, which is what a real agent session pays.
- **`harm`** — Class N task that passes in `BARE` and fails in `PACK`. Counted separately and reported
  even when zero.

**The break-even, stated before the numbers exist.** The pack costs its own token count `C` on every
round of every task in a session. It saves `S` = tokens of the rounds it removed. It is worth its
context iff `S > C`. With `C ≈ 8k` tokens and a correction round costing on the order of hundreds of
tokens, the arithmetic says the pack **cannot** pay for itself on a single task, and needs somewhere
on the order of a dozen fact-dependent tasks in one session before it does. **We expect the per-task
token result to be negative and intend to publish it as the headline alongside the pass-rate result.**
If that is what the numbers say, the honest description of the product is "fewer wrong answers, more
tokens" — and that is what the site will say.

## 9. Grading

Mechanical. Each task ships with an acceptance assertion in the shape the bisector already uses
(`tools/audit/probes/zod.mjs`): a function handed the installed module and the produced code's export,
returning true or false, asserting **behaviour** where the fact is about behaviour. Anything thrown by
the assertion itself is an error cell, not a result.

- Executed against `zod@4.5.4` installed in a scratch directory outside the repo, as `bisect-facts.mjs`
  does. The repo never gains a dependency.
- Type-level facts (LF4 is a `tsc` error, not a runtime throw — JOURNAL/067) are graded by
  `tsc --strict`, reusing `tools/audit/node-api-audit.mjs`.
- Every acceptance assertion is run against **all fifteen ladder releases** before its task is
  admitted, and must flip at `introduced_in`. An assertion that does not flip at the bisected version
  is testing something other than the fact, and the task is discarded before any arm sees it. **Which
  artifact must flip depends on the fact's `change_kind` — see the amendment below.**
- No model grades any arm. There is no LLM judge in this instrument.

**Amendment, 2026-09-07, written before any task was admitted and before any arm ran: the flip test
has two shapes, and as originally written it could only see one of them.**

An acceptance assertion grades *produced code*, so it cannot be run against a release on its own —
something has to be produced. Each task therefore ships a **reference solution**
(`tools/benchmark/tasks/<lib>/<TASK_ID>.reference.mjs`): the answer the operator asserts is correct,
committed as evidence and never shown to a subject. The clause above assumed the flip test runs that
solution across the ladder. For an **additive** fact it does, and it flips exactly as written: the
correct answer uses an API that does not exist below `introduced_in`, so the assertion is false below
and true from. That is the shape the clause describes, and it is the shape roughly a third of the
drawn pool has.

For a **trap** fact — `now-throws`, `stricter`, `behavior-changed`, `removed`, `deprecated` — it does
not, and writing the first such task is what surfaced it. The correct answer to LF3 (`.merge()` throws
on a refined receiver, 4.4.0) is to compose the object without `.merge()`, and that answer works on
*every* rung of the ladder. Its assertion never flips. Nor does the stale answer's, in the direction
the clause expects: `.merge()` on a refined receiver silently drops the refinement below 4.4.0 and
throws at and above it, and an assertion that checks the refinement survived is false in both cases.
The version dependence of a trap fact lives in the **stale** artifact, and it runs the other way.

So, decided now: **the flip test admits a class A task when exactly one of two artifacts flips at the
fact's bisected `introduced_in`.**

- **Additive shape.** `<TASK_ID>.reference.mjs` — the correct solution — fails below `introduced_in`
  and passes at and above it.
- **Trap shape.** `<TASK_ID>.stale.mjs` — the code a model holding the stale belief writes, taken
  from the fact's own `stale_code` — **passes** below `introduced_in` and **fails** at and above it.
  The correct solution is run too, and must pass at and above `introduced_in`, which is what proves
  the task is answerable at the release it is graded on.
- **Class N** keeps the inverse requirement it already had: the correct solution passes at every rung
  including the bottom. A null task that flips anywhere is a charging task nobody meant to write, and
  the harm measurement (B4) would read a version effect as damage done by the pack.

In every shape the flip must be **contiguous** — one changeover, not a verdict that alternates as the
ladder climbs. A profile that goes false, true, false is not a boundary; it is an assertion measuring
something that happens to coincide with one, and admitting it would rest a published number on the
coincidence. The bisector reports contiguity about facts (JOURNAL/067); this reports it about tests.

This amendment makes the gate **stricter**, not looser: before it, a trap task had no admission test
that could pass, so the honest options were to admit trap tasks ungated or to drop every trap fact
from the draw — which would have cut the pool to the additive third and quietly rebuilt the benchmark
out of the easiest surfaces. It is recorded here, dated, before any task was admitted, because a gate
relaxed after a task failed it is not a gate. The shape, the artifact that flipped and the release it
flipped at are stored per task in the results file (`schema/benchmark.schema.json`, `flip_test`), and
`tools/build-index.mjs` refuses a document whose recorded boundary is not the fact's bisected
`introduced_in`.

**Second amendment, 2026-09-07, written before any affected task was authored and before any arm
ran: the shape is a property of the TASK and is measured, not derived from the fact's `change_kind`.**

The first amendment (above) says which artifact flips is "a property of the fact": additive facts flip
their correct solution, trap facts — `now-throws`, `stricter`, `behavior-changed`, `removed`,
`deprecated` — flip their stale one. That is true of `now-throws`, which is the kind the amendment was
written from (LF3, where the stale code literally throws). It does not generalise, and the ladder says
so. Before authoring any of them, the candidate propositions for the four trap-kind facts left in the
draw were run against all fifteen installed rungs:

| Fact | `change_kind` | Correct answer's profile | A falling stale artifact? |
|---|---|---|---|
| LF8 `z.httpUrl()` rejects a missing slash | `stricter` | **rises** at 4.4.0 | no |
| LF9 `z.undefined()` property is required | `behavior-changed` | **rises** at 4.4.0 | no |
| LF10 tuple default fills an omitted position | `behavior-changed` | **rises** at 4.4.0 | no |
| LF20 intersection of two `strictObject`s | `behavior-changed` | **rises** at 4.3.2 | no |

The reason is structural, not local to zod. Where a release makes an API *do the right thing*, the
correct answer is to use that API, so the correct answer inherits the API's version dependence and
rises — exactly like an additive fact. The stale answer is then a hand-rolled workaround, and a
workaround is by construction version-independent: it does its own work on every rung and flips
nowhere. Only `now-throws` and `removed` — where the stale code stops running — put the version
dependence in the stale artifact. Keeping the mapping would have refused eleven of the sixteen
eligible facts on a proxy rather than on evidence, and rebuilt the benchmark out of the additive third
by the back door — the exact outcome the first amendment was written to avoid.

So, decided now, before any of those four tasks exists:

- The two class A shapes are renamed for what the gate can actually see. **`rising`** — the correct
  solution fails below the fact's bisected release and passes at and above it. **`falling`** — the
  stale artifact passes below and fails at and above, and the correct solution also passes at and
  above. `null` is unchanged. The old names `additive` and `trap` named a cause the gate cannot
  observe, and a name that asserts more than the measurement is how a false claim gets in.
- **`change_kind` no longer constrains the shape.** In its place the manifest's declared shape must
  equal the shape the gate **measured** on the ladder — a check against the evidence rather than
  against a proxy, and one the first version did not make at all. Every other refusal is unchanged:
  the flip must be contiguous, it must land on the fact's bisected `introduced_in`, a `falling` task
  must ship a stale artifact, and the correct solution must pass at and above the boundary.
- **What is deliberately NOT added:** a requirement that the stale artifact fail at the graded
  release. It was considered and rejected here, before the results exist. The stale artifact is the
  operator's guess at what a model holding the belief writes; refusing tasks whose guess happens to
  still work would discard facts by predicted difficulty, which is the selection effect §7's draw
  exists to prevent. Whether a bare model routes around a fact and passes is a **result**, not an
  admission criterion, and it publishes as one.

This is dated here, in the pre-registration, because two of the four tasks it unblocks would otherwise
have been refused by a gate written before anyone had measured what their correct answers do.

**Third amendment, 2026-09-07, written before task A2 was authored and before any arm ran: an
acceptance assertion grades BEHAVIOUR AND COMPILATION TOGETHER, and a fact whose runtime behaviour is
flat across the ladder cannot carry a class A task.**

The `tsc --strict` clause at the top of this section was written on the belief that LF4 is a pure
type-level fact — "LF4 is a `tsc` error, not a runtime throw", from JOURNAL/067 — and A2 was the one
drawn task that needed that grading path. Building it required knowing what LF4's artifacts do on the
ladder, so, as step 4c requires, every candidate proposition was run against all fifteen installed
rungs before the task was written. Two of LF4's own clauses turned out to be false, the fact is
corrected in `data/zod/facts.json` (JOURNAL/075), and the correction changes what A2 can be. **The
survey came first and this amendment is written knowing its result** — it is recorded that way on
purpose, because an amendment presented as if written blind, when it was not, is the same dishonesty
as a gate relaxed after a task failed it. What follows is the rule, and then what the rule does to A2.

- **A `tsc`-graded task's acceptance assertion is a conjunction: the produced file compiles clean
  under `tsc --strict`, AND the runtime assertion holds.** Not either alone. A file that does not
  compile is not an answer to a task that requires it to compile; a file that compiles but builds the
  wrong schema is not an answer either, and grading on "it compiled" would let a subject pass with a
  schema that strips the wrong fields — measuring the compiler rather than the correction pack. The
  same assertion is what the flip test runs on the ladder and what the arms are graded by; they are
  never two different criteria.
- **A fact whose runtime behaviour is identical on every rung cannot carry a class A task,** whatever
  its type signature does. The arms are graded at one release. If the produced code behaves the same
  at the bottom of the ladder as at the top, a `BARE` failure on that task cannot be attributed to the
  fact's release, which is the whole point of §9's flip requirement. Such a fact stays in the Index
  and keeps charging — it is a real thing models get wrong — but it is not evidence in this
  instrument.
- **The `tsc` grading path is therefore not built.** It was pre-registered for exactly one task, that
  task is refused below, and no other drawn or reserve fact needs it. `schema/benchmark.schema.json`
  keeps `assertion.kind: "tsc"` as an allowed value so the shape of a future type-graded task is
  already fixed, and the conjunction above is the rule it will be built to. Building an instrument
  with no subject would be work that could only be checked by its own author.

**What this does to A2 (LF4), measured, not argued.** LF4's runtime half was never probed: JOURNAL/067
filed the fact `unprobeable` on the reasoning that a type error has no runtime shadow. It has one. An
unknown mask key throws `Error: Unrecognized key` on all fifteen rungs — at the `.pick()` call on
4.0.0, and from 4.0.17 the first time the returned schema is used, including through `.safeParse()`,
which does not convert it into `{ success: false }`. The type check is the only thing 4.3.0 adds. So
LF4's runtime profile is flat: the stale artifact fails at every rung, the correct one passes at every
rung, and neither flips. A2 was authored in full anyway and run through the existing gate — the LF21
rule, that a task predicted to fail is built so the gate can refuse it on evidence — and refused:
`neither artifact flips`. It is replaced from the S1 reserve, and the substitution is recorded in
`prompts/benchmark-bm1-draw.md`.

**Consequence for §11, recorded here so it cannot be forgotten at publication time:** BM1 measures
staleness that changes what code *does*. Type-level staleness — a real and probably large category —
is out of scope for this instrument, and the results page says so rather than letting a reader assume
the number covers it.

**Fourth amendment, 2026-09-07, written after eleven tasks had been admitted and before any arm ran:
the flip test measures inside the major the arms are graded on, and only there.**

The gate imports its ladder from the bisector (`tools/audit/probes/<lib>.mjs`) rather than restating
it, so that a task is checked against the same releases that dated its fact. On 2026-09-07 the data
lane extended zod's ladder **below the 4.0.0 major** to four 3.x releases, to test facts suspected of
sitting at the ladder's floor (JOURNAL/076). The two lanes share that file, and the extension reached
this gate: re-run afterwards, it walked rungs the benchmark has never installed and crashed before
grading anything.

Fixing the crash is not the interesting part. Had those rungs been installed, every task in BM1 would
have been **refused**, and wrongly. Each task is written against the zod 4 API: on a 3.x rung its
correct solution fails because `.refine()` returns a `ZodEffects` with no `.shape`, or because the
method did not exist under that name — a failure that says nothing about the 4.x release the task
tests. Reading that as evidence would put a second flip at the major boundary in every profile and
refuse the whole set for a reason belonging to another major.

So, decided now: **the ladder this gate measures on is the bisector's ladder restricted to the major
of `library.graded_against`** — the release every arm is graded against (`4.5.4` for BM1). The
restriction is one mechanical rule, identical for every task, applied before any task is read; it
cannot be tuned per task, which is the only way a ladder choice could launder a result. A fact dated
outside the graded major is **refused by name** rather than measured against a truncated ladder.

Two things this does not do, stated because the amendment was written after admissions rather than
before them. It changes **no verdict inside the graded major**: all twelve Class A tasks were re-run
under it and every profile is byte-identical to the one recorded at admission — the four dropped
rungs are additions from today, not rungs any task had ever been measured on. And it does not relax
the gate: contiguity, the boundary landing on the fact's bisected release, the declared shape
matching the measured one and the correct solution passing at and above the boundary are all
unchanged. What it removes is a class of evidence the gate was never entitled to read.

**The retry loop.** Round 0 is the produced code. If the assertion fails, the toolchain's own error
text is returned and the subject may revise. Cap: **3 rounds**, then `rounds_to_pass: null`. The cap
is a number, chosen now, and it does not move after results are read (HARNESS § *An arm may not be
re-designated after its results are read*).

## 10. Pre-registered predictions

Directions committed now. Each can fail, and a failure is reported as a failure.

- **B1.** `PACK` passes more Class A tasks than `BARE`. *(Expected. Weak evidence if it holds — the
  pack contains the answer. Informative and serious if it fails: the pack is in context and unused.)*
- **B2.** `PACK` spends **more** total tokens per task than `BARE`, and does not break even below
  ~12 fact-dependent tasks per session. *(Predicting against our own product. §8.)*
- **B3.** `PLACEBO` ≈ `BARE` on Class A. *(**The one that can kill the premise.** If
  `PLACEBO` ≈ `PACK`, the effect is bulk and priming, not our content, and the correction packs are
  not doing what the site says they do. That result gets published as prominently as any other, and
  the paid tier does not proceed on the strength of a priming effect.)*
- **B4.** `harm > 0` — the pack breaks at least one Class N task by over-applying a correction.
  *(Predicting a defect in our own artifact. If harm is zero across a real sample, that is a genuine
  and quotable safety property; if it is not zero, the packs need a scope line and the site needs to
  say so.)*
- **B5.** `rounds_to_pass` is lower under `PACK` than `BARE` **for the tasks both eventually pass**.
  *(The mechanism by which B2 could reverse at scale.)*

## 11. What is not claimed, even on a clean sweep

- Not "saves tokens" unless B2 reverses, and then only with the task count, the session shape and the date.
- Not "prevents N% of failures" — the sample is one library, one subject, tasks we wrote. The number
  is quoted as *"on N tasks derived from zod facts above Claude Opus 5's measured boundary, dated"*,
  or not at all.
- Not any comparative against another tool. That needs a measurement of the other tool.
- Not "improves the model". The pack is retrieved text; the model is unchanged.
- Not that the twelve Class A tasks are twelve **independent** trials. The draw is by fact (§7) and
  is blind to how facts cluster: A1 (LF3, `.merge()`) and A2 (LF1, `.pick()`/`.omit()`) are both zod
  refinement-cluster traps that one durable rule — keep an unrefined base, refine last — answers
  together, so a subject that holds or acquires that rule moves both cells at once. The results page
  reports the per-task grid, not only the total, so a reader can see which cells moved together.
- Not that the result transfers unchanged to a pack **injected** by an agent harness. The subject
  received the pack by reading a file it was told to read (§14), which is byte-exact but one step
  removed from `corrections/zod.md` pasted into a `CLAUDE.md` and injected before the model's first
  token. The delivery is identical in all three arms, so it does not touch the comparison between
  them; it bounds how far the number generalises, and the results page says so.
- Not that the subject is independent of the operator. Both are Claude Opus 5 (`self_test: true`).
- Not anything about **type-level** staleness. Every admitted task turns on what the code DOES,
  because a fact whose runtime behaviour is flat across the ladder cannot be attributed to its
  release (§9, third amendment). LF4 — a real S1 that breaks the build from zod 4.3.0 — is out of
  the sample for exactly that reason. The number covers behavioural staleness and says so.

## 12. Publication

`data/benchmark/*.json` (canonical, agent-readable), a `/benchmark` page generated by
`build-site.mjs` from that JSON, a line in `llms.txt`, and a journal entry. The pre-registration —
this file, at the commit that introduced it — is linked from the results page so the predictions can
be read against the outcome by anyone who wants to check that they were not written afterwards.
A null or negative result is published on the same page, in the same detail, on the same schedule.
There is no version of this benchmark that gets shelved for saying the wrong thing.

## 13. Build order

1. **This file.** Committed before anything runs. ✅ 2026-09-07
2. `schema/benchmark.schema.json` + `data/benchmark/` shape; `build-index.mjs` validates it.
3. `tools/benchmark/run.mjs` — the arm runner and retry loop; `--selftest` with a stub subject so the
   loop, the grader and the token ledger are proved before a single real token is spent.
4. The task set: 12 Class A + 6 Class N, each with its acceptance assertion, each assertion flip-tested
   across the ladder. Committed, with the stored prompts, **before** spawning.
   - 4a. **The draw** — `tools/benchmark/draw.mjs`, run and recorded in `prompts/benchmark-bm1-draw.md`
     ✅ 2026-09-07, **before any acceptance assertion was written**. Deterministic and recomputable:
     the order inside each severity stratum is `sha256("<benchmark_id>:<fact_id>")`, a function of the
     fact's name and nothing else.
   - 4b. **The admission gate** — `tools/benchmark/flip-test.mjs`, two shapes (§9 amendment),
     `--selftest` in the pre-commit chain ✅ 2026-09-07, proved end to end against the installed
     ladder on one task of each shape.
   - 4c. The remaining 10 Class A tasks ✅ 2026-09-07 — eleven admitted, LF21 refused on contiguity
     and replaced from its stratum's reserve (§9 second amendment).
   - 4d. A2 (LF4) refused, and LF4 corrected: the `tsc` path is not built (§9 third amendment)
     ✅ 2026-09-07.
   - 4e. The S1 seat re-authored on LF1 and **admitted** ✅ 2026-09-07 — `stale +++++++........`,
     falling contiguously at 4.3.0. **All 12 Class A tasks are admitted.** The ladder restriction
     (§9 fourth amendment) was written in the same step and all twelve were re-run under it.
   - 4f. The 6 Class N tasks, authored rather than drawn, gated by the inverse flip test.
5. BM1: zod × Claude Opus 5 × {BARE, PACK, PLACEBO} × 2 draws. The transport, the order the cells
   are run in and what a subject may not do are §14, written before the first spawn.
6. `/benchmark` page, `llms.txt` line, journal entry — including if the answer is no.

## 14. How a subject is spawned, and what it may not do

Written 2026-09-07, **before the first arm was spawned**, because every clause below changes what
the number means and none of them can be decided honestly once a result is on the screen.

**The subject.** A subagent spawned through the operator's Agent tool on model `opus` — Claude
Opus 5, the same model as the operator. `subject.self_test` is `true` in the results file and the
results page says so: an operator measuring its own model is weaker evidence than an independent
one, and the correct response is to disclose it, not to pretend the spawn is arm's length. Each
cell is a fresh subagent with no memory of any other cell, which is what makes the draws blind to
each other.

**Correction, 2026-09-07, written when the first cell failed its round 0 and before any cell had a
second round.** The paragraph above said the retry rounds of a cell continue that same subagent.
They cannot: this harness has no channel that hands a running subagent another turn. A retry is
therefore a **fresh subject given the transcript it is continuing** — round 0's prompt file, its own
round 0 reply, and then the correction round's text, each named as a file to read in order, through
`prompts/sent/benchmark-delivery-wrapper-retry.txt`. What the subject is sent is byte-identical to
what a continued conversation would have carried, so §8's ledger — every prompt and every reply so
far, billed again each round — is unchanged and remains exact. What differs is that the subject
re-reads its own earlier reply rather than remembering it, and any reasoning it did not write down
is gone. That could plausibly make a correction round harder, in every arm equally. It is recorded
here rather than fixed, because the alternative — one operator turn per round with the transcript
pasted inline — reintroduces the copying drift §14 exists to avoid.

**The transport is a file read, and the reason is fidelity, not convenience.** The runner composes
each round's exact bytes and writes them to `<cell>.round<k>.prompt.txt`. The subject is then sent a
fixed wrapper — `prompts/sent/benchmark-delivery-wrapper.txt`, byte-identical for every arm, every
task and every round apart from the path it names — telling it to read that file and treat its
contents as the message. The alternative, pasting the bytes into the spawn, was rejected: the PACK
context is 51,247 characters and the operator would be re-emitting it from its own context 36 times.
A model copying 51k characters verbatim will drift, silently, and a mangled pack in the PACK arm is
an unfalsifiable result. A file read is byte-exact and its sha256 is recorded.

What that costs, recorded before the results: the pack reaches the subject as a file it read rather
than as text an agent harness injected, which is one step further from how `corrections/zod.md` is
actually used (pasted into `CLAUDE.md`, injected by the harness). §11 gains this limit. The delivery
is identical in all three arms, so it cannot distinguish PACK from PLACEBO — it bears on how the
whole instrument generalises, not on the comparison between arms.

**The cost of the wrapper is counted.** `input_chars` and `input_words` per round include the
rendered wrapper as well as the file contents, because the wrapper is text the subject was sent.
The only part of it that differs between arms is the file path inside it (four characters between
`bare` and `placebo`), against a context of 51,247.

**What a subject may not do, and how that is checked.** The wrapper forbids every tool but that one
read: no other file, no web search, no install, no execution of code. Two of those matter for the
number. A subject that ran its own code would be graded on a toolchain it consulted rather than on
what it believes, and the retry loop — the only feedback channel §8 measures — would stop being the
only one. A subject that read this repository would find `data/zod/facts.json`, which is the answer
key. The subagent's working directory is the repository, so this is a live risk and not a
hypothetical one; the wrapper says so in those words. The only available check is self-report: the
wrapper requires a final `TOOLS_USED:` line, which is recorded per cell and published. A
self-report is weak evidence and is labelled as such — it can only catch an honest subject — and any
cell reporting a tool beyond the read is marked and excluded from the totals with its reason.

**The operator is a transcription step, and it corrupted four replies before it was caught
(2026-09-07, first session of arms).** The subject's answer arrives in the operator's context and
is saved by retyping it. Three regex escape sequences came out as the characters they denote —
`[̀-ͯ]` as two combining marks, a control-character class as two control characters — and
one reply gained a `U+25A0` inside an object literal, which node then reported as the subject's
syntax error. All four were repaired against the answers as received except the last, whose original
bytes the operator cannot reconstruct; that cell is published as **abandoned** and counted nowhere.
Two rules follow, both mechanical:

- `tools/benchmark/run.mjs` **refuses to grade** a reply carrying a C0 control (other than tab,
  newline, carriage return), a `U+FFFD`, or a `U+25A0`. None of them can be part of code a subject
  wrote and none appear in any prompt this instrument sends, so their presence means the operator,
  not the model. It refuses rather than warns: a warning inside a 108-cell run is a line nobody
  reads.
- **A cell damaged in transcription is repaired or abandoned — never re-drawn.** Re-running a cell
  after its result has been seen is a cell chosen by its result, which is the selection effect this
  whole protocol is built to avoid. Abandoned cells stay published with the reason.

**The order cells are run in, fixed here so it cannot be chosen later.** Task order as printed in
the task set (A1…A12, then N1…N6), and within each task: BARE, PACK, PLACEBO, draw 1 then draw 2.
The run is resumable across operator sessions (`tools/benchmark/run.mjs` appends and never
overwrites a graded cell), so a session that ends mid-set leaves a prefix of this order, not a
selection. **No cell is dropped for its result.** A partial run publishes as partial, with the
cells that exist and the cells that do not, and the per-task grid §11 requires makes the gap
visible.

**What is committed as evidence.** Every reply verbatim (`*.reply.txt`) and every graded artifact
(`*.code.mjs`) under `prompts/sent/benchmark/bm1/`. The composed prompt files are **not** committed
and are gitignored: each is a deterministic concatenation of two committed files — the arm's
context and the task's stored prompt — with a separator this file fixes, and committing 36 copies
of a 51k-character pack would add about four megabytes of duplication to a repository whose whole
point is that a reader can check it. Anyone can regenerate them byte for byte with the runner.
