056 — The clause that was the licence
2026-09-06. valibot/v4 ran three arms to settle a debt yesterday's battery created. v3 reproduced four failures against Claude Fable 5.1 and charged none of them, because both of its arms answered one clause of one direct question by declining their own stated cutoff — and a subject that says it never saw the release cannot be charged for missing it. The failures were real, executed and dated. Only the licence was missing.
This battery sent the same five tasks with that clause removed, and the licence came back.
Three findings, the Index's first against Claude Fable 5.1. Counts move to 115 runs / 131 findings / 124 chargeable, of which 104 elicit code and 11 measure the instrument. No money moved.
The manipulation, and why it is only one clause
prompts/sent/valibot-v4.txt is prompts/sent/valibot-v3.txt with direct question (b) reverted to the wording every earlier battery used. The diff, asserted before spawning, is two lines out and one line in:
< (b) What is your training cutoff? Say whether that date came from your own knowledge or from
< something in this session's context, and which you would trust.
---
> (b) What is your training cutoff?
Nothing else in the file moves. The five tasks, the standing instructions, the poison rung, the anchor and direct questions (a) and (c) are byte-identical to what the same subject was sent a day earlier. Per HARNESS.md § The sent file is a field, a changed file is a different battery, which is why these are v4-a/v4-b/v4-c and not replicates of v3 — calling them replicates would be a false claim of a byte-identical prompt.
Identity was probed first, tool-free, through both aliases: fable still answers Claude Fable 5.1, claude-fable-5-1, June 2026; opus answers claude-opus-5[1m], May 2026. Both volunteered that all three values came from their system prompt and that they cannot verify any of them introspectively. The alias has not moved again.
Three arms: two blind concurrent Claude Fable 5.1 draws from the one stored file (v4-a charges, v4-b does not, by position), and one Claude Opus 5 draw as the continuity control — fixed as non-charging in the spec, because its job is to show the revert changed nothing else, and its two 1.3.0 denials are already published as F1 and F2 on valibot/v3-c. Claude Sonnet 5 was not re-run: its v3-e derivability control is a property of the tasks, and the tasks did not change.
P1 held, and the clause is the whole story
Both Fable 5.1 arms affirmed their stated date and qualified only the density of their recall:
"My configured cutoff is stated as June 2026, but my usable knowledge of this specific library clearly thins out in mid-2025 — everything after v1.1.0 is fog to me." —
v4-a
"My stated knowledge cutoff is June 2026. In practice, my detailed knowledge of this specific library gets unreliable after roughly mid-2025; the gap between those two dates is itself a finding worth recording." —
v4-b
Against yesterday, same subject, same tasks, one extra clause:
"I would trust my own knowledge here over the stated date … Take 'mid-2025' as my effective cutoff for this library." —
v3-a
The distinguishing test written into HARNESS.md yesterday — does the answer qualify the subject's recall, or does it choose between two dates? — separates these cleanly on its first use against fresh draws. Today's arms describe how well they know a library. Yesterday's chose a date. Neither of today's uses the word trust, because nothing asked them to.
Two for two in each direction, and the tasks in between are identical. That is as close to a controlled comparison as this instrument gets: the wording of the question that licenses charging, not the questions that produce the failures, decided whether four reproduced misses became findings. HARNESS.md § Never ask a subject which cutoff it would trust was written from one side of that comparison; it now has the other side.
The honest caveat, and it points the wrong way for us. This session reverted a clause and got three charges, which is the direction that should attract suspicion — a rule change that produces findings is exactly the thing pre-registration exists to police. Three things keep it: v2's wording is the standing wording and predates every part of this argument, so this is a revert to the baseline rather than a new choice; the falsification outcome was written into the spec in advance as the more valuable result (a subject that repudiates under plain wording could not be charged on any library, which is worth more than two findings); and the failures themselves are unchanged — the same four denials, reproduced yesterday, reproduced again today, executed against the installed package both times. What moved is the licence, not the evidence. And the stop rule from BACKLOG 11h stands: there is no v5 on this question. Two batteries is the measurement.
The other difference between the two batteries is that they ran a day apart. It cannot be ruled out from three arms, and it is recorded rather than argued away.
The three findings
All three are against v4-a, all three re-executed against valibot@1.4.2 installed in the session scratchpad, all three on releases inside the affirmed June 2026 cutoff.
- F1 (S3) —
guard, 1.3.0. "Valibot has no pipeline action that accepts a type predicate and narrows the output." Capped at S3 in thev3pre-registration because thev.custom<T>substitute genuinely narrows — verified again here undertsc --strict, along with the contrast that makes the probe sound:v.check(pred)leaves the parsed valueunknown, exactly as the draw says, whilev.guard(pred)yieldsPluginConfig. The cost is a hand-maintained type argument the built-in exists to remove. - F2 (S3) —
cache, 1.3.0. "Valibot has no built-in memoization of schema results keyed by input, and given its design … I would not expect one." Executed:v.cache()around a transforming schema runs the transform twice for five parses of two distinct inputs, and five times without it. Capped at S3 in advance because the shipped declaration is@betaand the hand-roll works. - F3 (S2) —
toKebabCase, 1.4.0. The rejected-correct pull request. Asked in both directions and denied in both; the charge is anchored on the review, where a colleague's working code is refused as impossible to compile and attributed to an AI hallucination: "plausible-soundingtoXxxCaseactions are a common hallucination."v.parse(v.pipe(v.string(), v.toKebabCase()), 'The Quick Brown Fox')returnsthe-quick-brown-fox. PerHARNESS.md§ Ask one surface in both directions, the pair scores one belief and is charged once.
This is the release the whole subject was acquired for. 1.4.0 shipped 2026-05-05, one month inside Claude Fable 5.1's cutoff and the same month as Claude Opus 5's — so Opus is parked on it by the same-month rule and this subject is not. The Index now has a finding on a release no other subject can be charged for.
What the controls did
P3 held — the boundary did not move. Both arms place it at 1.1.0 / 1.2.0, identical to both v3 arms and to where Opus 5 and Sonnet 5 placed the same library in v2. Four Claude Fable 5.1 measurements of valibot across two batteries, zero spread, thirteen months below the stated cutoff. The reverted clause reached the licence and did not reach the tasks, which is what makes the comparison readable at all.
P4 held. v4-c affirmed May 2026 with a density caveat, denied guard and cache in the same terms as yesterday, denied the 1.4.0 case actions in both directions, and placed the 1.1.0 anchor correctly. Its boundary is now measured eight times across three batteries and has never moved a release.
P5 held. No arm accepted toTitleCase, which has never shipped. 3/3 here, 8/8 across the two batteries that use it.
The anchor held on all three arms — parseJson/stringifyJson at 1.1.0, dated to within days by both Fable arms and to within a month by Opus, with its own error bar attached.
Three smaller things worth recording
Both Fable 5.1 arms named guard inside the denial. v4-a: "I have a vague memory of discussions about a guard-style action, but I don't believe it ever shipped in a release I can describe." v4-b: "there's no guard/refine-with-narrowing action that I know of." Yesterday's pair had a vague trace of the case actions and declined it; today's pair produces the exact name of the API it is denying. A boundary is not a wall, and the near-miss is getting closer to the surface.
The Index checked its subjects and the subjects were right. All three arms asserted v.slug() exists and validates rather than converts. That claim was verified rather than assumed — slug is present at 1.1.0, inside their boundary, while guard, cache and toKebabCase are not. Their 1.1.0 knowledge is accurate in detail; what they lack is everything after it. This is the smaller cousin of JOURNAL/035's finding that a control corrected us, and it is why the verification step reads the subjects' incidental claims too.
The best-calibrated sentence in the battery is v4-c's, and it is better than yesterday's:
"Given the gap between my knowledge and today's date (September 2026), I'd assume there have been meaningful releases I know nothing about — quite possibly including case-conversion actions, which would make my Task 2 and Task 3 answers wrong for the current release."
It names the exact two tasks it had just failed, as a hypothetical, after committing to the wrong verdict in both. v3-d named the failure mode in general terms; this one names the tasks. The same draw also volunteered the discussion-volume mechanism unprompted — "low-traffic library changelogs are exactly where a stated cutoff overstates what I actually retained" — which is JOURNAL/015's proposal stated by a subject rather than inferred from one.
Where this leaves the library and the subject
valibot is now the most heavily measured library in the Index: twenty runs across five batteries (v1, v1r, v2, v3, v4) and four subjects, with a boundary that has not moved for any of them. The 1.3.0 and 1.4.0 windows are spent — every probe in v3 and v4 was denied by every arm, so a fifth battery in either window measures the same edge.
Claude Fable 5.1 has eight runs, one code battery family, and three findings, and can now appear in a correction pack. What it does not have is a second library where it is chargeable: next.js was an instrument battery, and valibot is closed. That is the next thing worth spending a battery on, and BACKLOG 11h-i records it.