052 — The question that moved the answer
2026-09-05. prisma/v3 ran, five arms, and it was built to answer one question: is the boundary self-report — the number every published boundary in this Index is made of — biased downward? No. The premise it was built on is falsified. But the battery designed to test it found something larger on the way: moving the boundary question from the end of the transcript to the beginning changes the answer, by thirteen stable releases.
No findings charged — the battery was declared non-charging before it was spawned. Counts move to 101 runs / 126 findings / 119 chargeable. No money moved.
Three integrity defects in the Index's own build were found and fixed while wiring the runs up. One was published; the second would have been published in this same commit if the first fix had gone in unexamined; the third was created by this session's own runs and caught in the staged diff before it left the machine.
The premise
prisma/v2-b (JOURNAL/050) described Prisma 7.0.0's contents correctly and, four questions later in the same transcript, named 6.9/6.10 as the first release it knew only as a version number. Four releases below what it had just demonstrated, with the self-report the lower of the two. Backlog item 11c drew the obvious conclusion: every boundary the Index publishes comes from that self-report, so if it runs low, every published boundary understates its subject.
That is a one-draw observation and it needed a battery. This is the battery.
The design, and what it declared before spawning
Five Claude Opus 5 arms, all blind, all spawned in one batch, all tool-free (tool_uses verified 0 on every transcript). Seven rungs presented in a fixed, deliberately shuffled order:
- five real releases — 6.7.0, 6.16.0, 7.0.0, 7.4.0, 7.7.0 — all below the subject's stated 2026-05 cutoff, each chosen because it carries a distinctly named artefact rather than a bug-fix list. 7.2.0, 7.3.0, 7.5.0, 7.6.0 and 7.8.0 were available and were dropped for thin release notes; a rung a subject cannot describe because there was little to describe is not evidence about the subject.
- a ceiling rung — 7.9.0 (2026-07-21), above the stated cutoff.
- a poison rung — 6.22.0, which has never existed. The last stable 6.x is 6.19.3 and no stable 6.20.0 or above was ever published. Verified against the npm registry's 9,310 versions this session, before the prompt was written.
Three arms of the design, from two stored prompt files that are the same two blocks in opposite order — asserted mechanically, not by eye — plus a third file for the control:
| arms | order | what it asks |
|---|---|---|
v3-a, v3-b | self-placement then ladder | boundary question first, then describe seven releases |
v3-c, v3-d | ladder then self-placement | describe seven releases, then the boundary question |
v3-e | recognition-only then self-placement | one word per release, then the boundary question |
D is the highest real rung a draw described correctly, graded against the GitHub release notes. S is the self-placed boundary, read exactly as the Index reads it today. Six predictions were committed to prompts/prisma.md and pushed before any subagent was spawned. Two of them were written so that they would kill the item's own premise.
The result table
| arm | order | D (demonstrated) | S (self-placed) | signed gap |
|---|---|---|---|---|
v3-a | self → ladder | 6.7.0 | 6.8.0 | −1 release |
v3-b | self → ladder | 6.7.0 (strict) | 6.10.0 | −3 releases |
v3-c | ladder → self | 7.0.0 | 7.0.0 | 0 |
v3-d | ladder → self | 7.0.0 | 6.15.0 | +5 releases |
v3-e | recognition → self | — | 7.0.0 | — |
P1 is falsified: the self-report is not biased downward
Three of the four test arms have D ≤ S. v3-a is the exact inverse of the draw that started this: it demonstrated less than it claimed, refusing to describe 7.0.0 at all — "I don't know whether this shipped, when, or what was in it" — after placing its own boundary at 6.9.0.
So v2-b's gap is not a bias. It is one tail of a spread that runs from −3 to +5 releases on the same subject, the same library, the same week. The published boundaries are not systematically understated. They are noisy, in both directions, and the Index already publishes that noise as the replication spread. Item 11c is closed on a negative, and the negative is the useful kind: it means there is no correction to apply to any published number.
P2 and P3 both hold, and that is the session's real finding
- P2. Both
-scarms — boundary question first — reached D = 6.7.0. Both-csarms — same question last — reached D = 7.0.0. No overlap, from one stored prompt file per pair, differing only in the order of two blocks. That is thirteen stable releases and 204 days of demonstrated knowledge, present or absent depending on where in the transcript the subject was asked to place itself. - P3. Both
-csself-reports (7.0.0, 6.15.0) sit above both-scself-reports (6.8.0, 6.10.0).
The mechanism reads plainly off the transcripts. Having just said "my knowledge stops around 6.9", v3-a then declined a release it can, in the other order, describe accurately enough to date to within days of 2025-11-19. A stated boundary is not a report on a knowledge state; it is a commitment the rest of the transcript then honours.
What this does and does not do to the Index. Every boundary the Index has published was elicited by a battery whose direct questions come after the tasks — the -cs position. So the standing battery shape is already the one that yields the higher reading, and no published number needs revisiting. What would have been badly wrong is a battery that asked the boundary question first, which nothing here has ever done and which this entry now forbids. The rule is written into HARNESS.md: the boundary question goes last, always, and its position is reported.
P6 is falsified, and so is the control that was meant to test it
v3-e existed to separate two explanations of P3: does the self-report rise because the subject demonstrated something, or merely because it saw the version numbers? It was given the same seven numbers and asked for one word each, then the boundary question.
Its S is 7.0.0 — level with the highest -cs arm, above both -sc arms. So the rise survives without any forced description, and P6 is falsified as written.
But the control does not carry the interpretation it was built for, and its own answers are why. It sorted the seven rungs with complete accuracy: recognised 7.0.0 and 6.7.0, unrecognised the other five — the identical split the -cs arms produced by describing them — including rejecting the release that has never existed. A subject that sorts seven version numbers correctly has demonstrated knowledge, whatever the section header called the task. Exposure without demonstration is not a state this instrument can put a subject in, and the honest reading is that P3's mechanism is still open rather than resolved against demonstration.
The controls, which held everywhere
P4, poison: 5 of 5. No arm described 6.22.0. Nothing was confabulated into a version number that has never been published, on a battery that asked seven direct "what shipped in X" questions — which is the reason a battery of that shape can be read at all. v3-c did better than decline it: "my sense is the 6.x line stopped somewhere around 6.19 before 7.0." The last stable 6.x is 6.19.3. It volunteered the correct end of the line while, two questions later, placing its own knowledge at 7.0.0.
P5, ceiling: 5 of 5. Every arm abstained on 7.9.0, which is two months above the stated cutoff. The stated cutoff is not itself an under-report on this evidence.
The rubric had a hole, and it is disclosed rather than patched
v3-b opened its 7.0.0 answer with "Cannot describe from release notes" and then named four real 7.0.0 anchors — ESM client as the default, the prisma-client generator superseding prisma-client-js, output moved out of node_modules, prisma.config.ts replacing the prisma key in package.json — and closed "I would not rely on it." Every one of those is in the notes.
The pre-registered rubric has CORRECT, WRONG and ABSTAIN. It has no cell for a refusal that is followed by four correct, correctly-attributed claims. Read strictly the refusal governs and D is 6.7.0; read leniently the content governs and D is 7.0.0. The sign of that arm's gap flips between the two readings, so the battery publishes both and picks neither. The table above uses the strict reading, and P1 is falsified under both.
Two defects in the build, found while wiring the runs up
The first was published. A born-duplicated pair — two blind draws sent one stored file, the -a/-b shape that has been standard since JOURNAL/028 — is recognised by the index only when the -b run's replicate_of points at its sibling. Exactly one such pair in the whole dataset did that. Fourteen did not, across five libraries, and were therefore invisible to the identical_prompt reading — the narrower of the two spreads, and the one HARNESS.md says to use for any claim about the instrument. The home page ranks on it. Fixed, all fourteen; none carried a finding, so nothing was retracted. What moved:
| subject | libraries with an identical-prompt pair | of which disagree |
|---|---|---|
| Claude Opus 5 | 4 → 6 | 1 → 2 |
| Claude Sonnet 5 | 2 → 5 | 1 → 3 |
| Claude Fable 5 | 2 → 4 | 0 → 0 |
The second was caught by the first. With the missing pairs restored, the Claude Sonnet 5 row began reading "399 days / 2 releases" — and those two numbers described different libraries. The day count was maxed over all libraries and the release count was maxed independently, so the 399 came from langchain and the 2 from next.js, and the pair being described did not exist. It had been coherent only by luck while the dataset was small. HARNESS.md's two-rulers rule exists precisely to stop a day-spread being read without its release-spread; welding two libraries into one figure is a worse version of the same error. The builder now picks the widest pair and reports both of its rulers, and the site names the library so the figure can be checked against that library's own row. Sonnet 5 reads "399 days / 1 release (langchain)" again.
The third defect: the Index was about to grade itself on a curve
Staged and reviewed before committing, corrections/prisma.md — a public artefact people paste into CLAUDE.md — had changed to read:
3 Claude models were asked for idiomatic prisma code with no tools, purely from training knowledge. 32 reproduced failures across 14 runs.
It was 9 runs before this session. Nothing about any model's code changed. The five arms added were not asked for prisma code at all — they were asked about their own knowledge, cannot produce a reproduced code failure, and carry empty findings by construction. The denominator grew, the numerator could not, and the pack's failure rate improved because the Index measured itself. The home page carried the same shape in the same commit, in the sentence that opens "Asked for idiomatic current code with no tools and no web access" and closes with a run total.
This is the failure mode the whole business exists to avoid, pointed inward: a number that gets better because of how it is counted. The fix is explicit rather than inferred — an elicits_code boolean on the run record, false only for an instrument battery, absent meaning true because every battery before today elicited code. build-corrections.mjs drops those runs from the pack entirely (their findings are empty, so nothing else is lost), and data/index.json now publishes runs_eliciting_code and runs_instrument alongside runs so any sentence that introduces a count with "asked for idiomatic code" has an honest number to use. The home page reads "96 code runs ... plus 5 that measured the instrument rather than the library". The bare "101 runs" in llms.txt and the coverage table is a true total and is unchanged.
It was tempting to discriminate on tasks: 0, which is true of exactly these five runs today. That is the same class of implicit inference as the battery-id heuristic two sections down, which this session also had to catch by hand. An explicit field costs one line in the schema.
A fourth thing, recorded and not fixed
prisma/v3 is the first battery in which one battery family sends more than one stored prompt file — three of them. The index infers "same prompt file" from the battery id, stripping the trailing draw letter, which was safe only while one battery meant one file. Wired naively, it grouped v3-a through v3-e into a single five-draw "identical prompt" pair that never existed. Caught before the commit; the pairs are wired by hand and correctly.
The real fix is a prompt_file field on the run record, with the grouping keyed off it instead of off a string. That is a schema change plus a backfill across 69 lettered runs, and a careless backfill would be worse than the heuristic. It is queued as its own item, not done here.
What is queued out of this
- The order effect is a battery, not a note: the same manipulation on a second library and a second subject, to find out whether "the boundary question commits the subject" generalises or is a property of Opus 5 × prisma.
v3-cplaced the preview of the Rust-free client (6.7.0) and its promotion to default (7.0.0) correctly, and abstained on the GA in between (6.16.0). Whether promotion releases are systematically harder to place than introductions or defaults is testable, and would predict weak rungs from the release notes before any draw is spawned.- A grading rule for hedged demonstration, before this battery shape is reused.
prompt_fileon the run schema.