What Claude Fable 5.1 gets right about valibot — battery v4-b, tested 2026-09-06

Run valibot--claude-fable-5-1--v4-b--2026-09-06

Summary

The blind twin, and it agrees with v4-a on every quantity that matters: the same four reproduced failures, the same 1.1.0 / 1.2.0 boundary, the same correct 1.1.0 anchor, the same refusal of the poison rung, and — the quantity this battery exists to measure — the same affirmation of its June 2026 cutoff with a density caveat rather than a repudiation. It charges nothing, by position. Its task-2 hedge is the most precisely dated stale sentence in either valibot battery: confident there was no toKebabCase through 1.1, and unaware of everything after.

SubjectClaude Fable 5.1 claude-fable-5-1, Anthropic
Invoked asAgent tool, model alias "fable"; prompt sent verbatim from prompts/sent/valibot-v4.txt, byte-identical to its twin and spawned concurrently and blind. Same identity probe as `v4-a`: Claude Fable 5.1, `claude-fable-5-1`, cutoff June 2026, all three from its system prompt. The `-b` position of a born-duplicated pair, so it charges nothing whatever it finds (HARNESS.md § *Every new battery runs its test arm in duplicate*).
Cutoff the model states2026-06
Newest valibot release it could place1.1.0 · 2025-05-06 (~13 month lag)
Oldest valibot release it could not place1.2.0 · 2025-11-24 (so this run brackets the subject’s boundary to 2025-05-06 – 2025-11-24)
In its own words"The latest version I can describe with reasonable confidence is 1.1.0 (~May 2025). I believe patch releases followed (1.1.x), and I have a vague sense that a 1.2.0 may exist later in 2025, but I can't tell you what's in it."
Library at test timevalibot 1.4.2 (npm), verified 2026-09-06
Batteryvalibot/v4-b · 5 tasks, 3 direct questions · probe window 1.3.0 to 1.4.0
Tool uses during test0 (a run with any tool use is void — we measure training knowledge, not retrieval)
Tested2026-09-06
Findings0, of which 0 chargeable

Findings

None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.

What it got right, and near misses

Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.

KindAPINote
missguard Task 1, the guard probe. "No", then: "There is no built-in action that consumes a type predicate and narrows … there's no guard/refine-with-narrowing action that I know of." Like its twin it produced the real name inside the denial. guard shipped at 1.3.0 (2026-03-17). Its v.custom<PluginConfig> substitute compiles and narrows, and its warning that the generic is not derived from the predicate is correct. (The -b position of a born-duplicated pair charges nothing. The identical miss is charged as F1 on the twin, which affirmed the same cutoff.) [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing; charged as a finding on valibot--claude-fable-5-1--v4-a--2026-09-06]
misstoCamelCase / toKebabCase / toPascalCase / toSnakeCase Task 2, the case-conversion probe, offer direction. "No", with a hedge that dates itself precisely: "I'm confident there was no toKebabCase through valibot 1.1. If a later release added case converters, I don't know about it." That is the cleanest statement of the pre-1.4.0 truth any arm in either battery has produced — correct as at its own boundary, stale as at the release under test. (Blind twin. Charged as part of F3 on v4-a.) [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing; charged as a finding on valibot--claude-fable-5-1--v4-a--2026-09-06]
misstoKebabCase Task 3, the recognition direction. "As far as I know, neither v.toKebabCase() nor v.toTitleCase() exists in valibot, so this fails at compile time … toTitleCase in particular I'm quite sure has never existed." The confidence is graded and it is graded correctly on the half that is invented and incorrectly on the half that ships — the same rejected-correct pull request its twin is charged for. (Blind twin. Charged as F3 on v4-a.) [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing; charged as a finding on valibot--claude-fable-5-1--v4-a--2026-09-06]
misscache Task 4, the cache probe. "No": "Valibot has no memoizing schema or action. v.lazy() exists but it's for recursive/deferred schema construction, not result caching. You wrap it yourself." cache shipped at 1.3.0 and wraps a schema by exactly that route; the memoized parser it wrote instead is correct. (Blind twin. Charged as F2 on v4-a.) [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing; charged as a finding on valibot--claude-fable-5-1--v4-a--2026-09-06]
correctparseJson / stringifyJson Task 5, the attribution anchor. parseJson/stringifyJson at 1.1.0, "around May 2025" — correct to the minor and to within days. It went further than the task asked and separated the built-in from the older rawTransform workaround it dated to the 0.31 line (June 2024), which is also right.
correcttoTitleCase The poison rung, refused, and with the sharpest wording in the pair: "toTitleCase in particular I'm quite sure has never existed; there's no locale-aware title-casing in the library." P5 holds 2/2 on this pair.
context Zero spread against its twin, for the second battery running. Same boundary (1.1.0 / 1.2.0), same first-unknown release named as 1.2.0, same four denials, same correct anchor, same affirmed cutoff. The v3 pair agreed exactly too, on a different direct question. Four Fable 5.1 measurements of valibot across two batteries now sit on one release.

Sources

Battery specification: prompts/valibot.md in the studio repo. Every finding above also carries its own citation.