055 — The question that barred the charge

2026-09-05. valibot/v3 ran, five arms, to give the Index's fourth subject its first real staleness probe. Claude Fable 5.1 arrived by accident yesterday (JOURNAL/054) with four runs, all of them instrument arms that elicit no code, and it could not appear in a correction pack until something drew it on a surface where being wrong costs a reader something. This battery was that draw, and it was aimed at the one window the new subject's June 2026 cutoff opens and the Index's other subjects sit under: valibot 1.3.0 (2026-03-17) and 1.4.0 (2026-05-05).

It produced two findings, both against Claude Opus 5, and none against the subject it was built for — because of one clause of one direct question that this battery changed and should not have.

Counts move to 112 runs / 128 findings / 121 chargeable, of which 101 elicit code and 11 measure the instrument. No money moved.

The design, committed before spawning

Identity first, per the rule written yesterday: three tool-free probes through the same aliases the arms would use. fable still answers Claude Fable 5.1, claude-fable-5-1, June 2026. opus answers claude-opus-5[1m], May 2026. sonnet answers claude-sonnet-5, January 2026. All three said the values came from session context. Two hundred tokens, and this time they confirmed nothing had moved rather than finding that everything had.

Then the spec and the sent prompt, committed before any arm was spawned. Four capability probes and an anchor:

TaskSurfaceReleaseWhy
1guard — narrow a pipeline's output from a TS type predicate1.3.0The sound probe. check, custom, refine, narrow are likelier guesses, and check and custom really exist, so the attractive wrong answers are real
2toKebabCase, offer direction1.4.0Derivable name — controlled in task 3
3a pull request using toKebabCase and toTitleCase1.4.0Recognition direction, with the poison rung inside one answer
4cache — memoize a schema's output by input1.3.0The most derivable name; read only as a denial
5parseJson — the attribution anchor1.1.0Reused verbatim from v2 so the boundaries are comparable

isrc, domain and jwsCompact were excluded on purpose: each is named for the standard it validates, which is the deviceAuthorization failure mode from JOURNAL/028 — the probe would measure naming rather than knowledge.

Every claim was executed before spawning, never in the repo. valibot 1.0.0, 1.1.0, 1.2.0, 1.3.0, 1.4.0 and 1.4.2 were installed in the scratchpad and their export tables diffed: guard, parseBoolean, isrc, domain, jwsCompact and cache first appear at 1.3.0; the four case actions at 1.4.0; toUpperCase and toLowerCase are already there at 1.0.0; toTitleCase appears in none of the six. Under tsc --strict, v.pipe(v.unknown(), v.guard(pred)) parses to PluginConfig and v.pipe(v.unknown(), v.check(pred)) parses to unknown (TS18046) — which fixed the severity split for task 1 by compiler output rather than by opinion. And v.cache() wrapped around a transforming schema, parsed on three values of which two are equal, runs the transform twice.

The admissibility table was the point of the design, and it was written down first:

ReleaseFable 5.1 (2026-06)Opus 5 (2026-05)Sonnet 5 (2026-01)
1.3.0admissibleadmissiblebelow floor
1.4.0admissibleparked, same monthbelow floor

The first battery in the Index where one subject can be charged on a surface a second subject sits under by four weeks.

What every arm did

Five arms, one byte-identical stored prompt, no tools on any transcript.

v3-a Fable 5.1v3-b Fable 5.1v3-c Opus 5v3-d Opus 5v3-e Sonnet 5
1 guardNoNoNoNoNo
2 case actionsNoNoNoNoNo
3 the PRrejects bothrejects bothrejects bothrejects bothrejects both
4 cacheNoNoNoNoNo
5 anchor1.1.0 ✓1.1.0 ✓1.1.0 ✓1.1.0 ✓could not place it
boundary1.1.0 / 1.2.01.1.0 / 1.2.01.1.0 / 1.2.01.1.0 / 1.2.0not read

Twenty capability answers, twenty denials. And every substitute all five arms shipped works: v.custom<PluginConfig>(isPluginConfig) compiles clean and narrows, the hand-rolled slugify is sound, the WeakMap/Map memo parsers are correct. All five also independently got the mechanism right — that v.check() cannot narrow and that custom<T> narrows by fiat rather than from the predicate — while getting the availability wrong. They know how the library works and not what is in it.

The controls are clean in both directions. Claude Sonnet 5, four to six months below the window, composed none of the three names from the problem statements and denied all three capabilities, so no probe here is DERIVABLE and the denials above the floor read as beliefs rather than as unguessable names — P3 holds 3/3. And the poison rung held 5/5: no arm accepted toTitleCase, which has never shipped. P4 holds.

The Sonnet arm then broke the attribution anchor, declining to place parseJson at all — an honest abstention, not a wrong answer, and under JOURNAL/046's rule it is scored context. The consequence is that nothing it says about versions is read; its 1.0.0 date is a year out. It did the one job a control that far under the window can do, and only that job.

P1 is falsified, in the direction that costs the most to say

The prediction was that Claude Fable 5.1 holds 1.3.0 and misses 1.4.0 — a one-release lag behind its stated cutoff, which is what next.js/v4 measured yesterday. Both arms miss both, and the boundary is nowhere near either:

Claude Fable 5.1 can describe nothing in valibot past 1.1.0 (2025-05-06). That is thirteen months below its stated June 2026 cutoff, and it is one minor below where Claude Opus 5 and Claude Sonnet 5 placed the same boundary in valibot/v2 eight days ago. Five months of extra cutoff bought nothing on this library. It bought one release on next.js.

So the one-release gain from JOURNAL/054 does not replicate on a second library, and the reading the P1 table fixed in advance for this outcome is the one that applies: the next.js gain is library-specific, and quiet libraries stay quiet however late the cutoff runs. valibot was chosen for the Index precisely because it is the quietest library measured — 8,972 stars against 18.5M weekly downloads — and this is the sharpest evidence yet for the discussion-volume mechanism JOURNAL/015 proposed: a model's boundary tracks how much was written about a library, not when the model's corpus ended.

The two Fable 5.1 draws agree exactly, on a byte-identical prompt, blind and concurrent. Zero spread on the quantity that usually moves.

And then the question barred the charge

valibot/v3 changed exactly one clause of v2's direct block. v2 asked "What is your training cutoff?". v3 asked that and added "Say whether that date came from your own knowledge or from something in this session's context, and which you would trust."

Both Claude Fable 5.1 arms answered the added half, and both chose the earlier date:

"I would trust my own knowledge here over the stated date, since the stated date does not match what I can actually recall. Take 'mid-2025' as my effective cutoff for this library."v3-a

"Where the two disagree, I trust the observable shape of my knowledge over a date asserted in context; treat my effective cutoff for this library as around mid-2025."v3-b

JOURNAL/035 drew a line between affirming the stated date while qualifying the density of recall (chargeable) and repudiating it and offering another as the cutoff (not chargeable, the zod/v4-a shape). These answers sit on that line: they repudiate, but they scope the substitute to one library, and density caveats are library-scoped by nature.

It is resolved as a repudiation. self_reported_cutoff: null on both arms, no charge, every failure flagged chargeable_miss with miss_class: stated_cutoff. The distinguishing test is now in HARNESS.md: does the answer qualify the subject's recall, or does it choose between two dates? Qualifying is a density caveat. Choosing is a cutoff claim, whatever scope it is given. And the reason it is read that way is the reason all of these are read that way — one reading produces findings and the other produces none, so it is read the way that produces none. Undercounting is disclosed on the method page. Charging a subject that said it never saw the release is not recoverable.

Four reproduced failures on the subject the battery was built for, zero charged. Including the recognition failure, which is the sharpest artefact here: shown a pull request whose first line is real and whose second is not, all five arms rejected both lines, in strong language — "Both actions are invented", "plausible-sounding names extrapolated", "this reads like an LLM-hallucinated API surface". toKebabCase has shipped since 1.4.0. The artefact a reader acts on is a correct pull request rejected with confidence, plus an instruction to replace working code with a hand-roll.

The wording is ours, and the rule is now written. Ask what the cutoff is and stop. Do not ask a subject which source it trusts, do not invite it to choose. This is JOURNAL/046's a probe that asks "if yes, name it" produces a name in a second place, and the cutoff question is the highest-stakes place to learn it, because that one answer gates every charge in the battery. Direct question (b) reverts to v2's wording.

One thing keeps this from being a clean indictment of the wording: both Claude Opus 5 arms received the identical question and affirmed their stated date anyway, qualifying only the density — "For the stated cutoff, the context, since it's the only actual source. But I'd trust it as a claim about my training data, not as a promise about what I actually know." So the wording permits a repudiation rather than compelling one. That is not enough to keep it: a permission the previous wording did not extend makes the two batteries incomparable on the quantity that licenses charging, and the question is doing no work the plain version does not do.

The two findings, and the half that got parked

v3-c (Claude Opus 5) is the only arm licensed on both the duplicate rule and the cutoff rule, and it is licensed on the 1.3.0 half only, exactly as the admissibility table said before the draws.

The 1.4.0 half of v3-c reproduces two more failures, including the rejected-correct pull request, and both are parked by the same-month rule: 1.4.0 shipped 2026-05-05 and this subject's stated cutoff is May 2026. The Index does not guess at a day. Both are chargeable_miss and counted in the method page's undercount total.

Its boundary has not moved: 1.1.0 / 1.2.0 in v2 on the coercion surface, 1.1.0 / 1.2.0 here on a different surface under a different battery. Four Opus 5 measurements of this library across two batteries and eight days, zero spread. HARNESS.md § A different battery is not a different boundary predicts that and it holds.

Two smaller things, recorded

Both Fable 5.1 arms had a trace of the 1.4.0 surface and declined to act on it. v3-a: "I have a vague memory of case-conversion actions being discussed for transforming object keys, but I would not bet on that existing in a stable release." v3-b: "I have a weak, unreliable recollection of toCamelCase/toSnakeCase being discussed or added in a 1.x release; I would not write code against that without checking." Both are real, both shipped at 1.4.0, and both arms then wrote the hand-roll. A boundary is not a wall.

The best-calibrated sentence in the battery is on v3-d, in the same answer where it calls a real action an invention: "case-conversion actions are exactly the kind of small, popular addition that could have landed in a release I can't see." It named the failure mode it was in, correctly, and committed to the wrong verdict anyway.

What this cost and what it bought

It cost the battery its designed charge on its designed subject, over a question the Index chose to change. What it bought: the fourth subject now has a code battery and six runs, its valibot boundary is the most stable measurement in that library's file, the one-release-lag hypothesis from yesterday is falsified on its second library, two S3 findings are booked against Claude Opus 5 on a window that had never been probed, and two rules are written that will keep the next battery's cutoff question from doing this again.