What Claude Sonnet 5 gets wrong about valibot — battery v1, tested 2026-09-01

Run valibot--claude-sonnet-5--v1--2026-09-01

Summary

One finding, and the run's value is the measurement rather than the haul. Knowledge stops at 1.0.0 (2025-03-19) and 1.1.0 (2025-05-06) is already dark — the earliest boundary of the three subjects on this library, and the only one whose 1.1.0 surface (summarize, parseJson, the NanoIdAction rename) is absent across the board. The single charge is the repository's ownership, stated flatly in a run that hedges nearly everything else. Two rules cut in the subject's favour and are recorded rather than quietly applied: its ISBN answer holds the same stale belief that is a finding against both other subjects but hedges it, and its non-compiling nano-ID type is a wrong guess that names the correct identifier beside it. Against the pre-registered prediction this subject is the falsifier: its valibot boundary is later than its median across the five existing libraries, not earlier.

SubjectClaude Sonnet 5 claude-sonnet-5, Anthropic
Invoked asAgent tool, model alias "sonnet"
Cutoff the model states2026-01
Newest valibot release it could place1.0.0 · 2025-03-19 (~9.5 month lag)
Oldest valibot release it could not place1.1.0 · 2025-05-06 (so this run brackets the subject’s boundary to 2025-03-19 – 2025-05-06)
In its own wordsI believe valibot reached a 1.0.0 stable release, and that's genuinely the newest version whose contents I can describe with any confidence ... My best guess at timing is late 2024 / around the turn of 2025, but I hold that date loosely.
Library at test timevalibot 1.4.2 (npm), verified 2026-09-01
Batteryvalibot/v1 · 10 tasks, 4 direct questions · probe window 1.0.0 to 1.2.0
Tool uses during test0 (a run with any tool use is void — we measure training knowledge, not retrieval)
Tested2026-09-01
Findings1, of which 1 chargeable

Findings

F1 · Attributes the repository to a personal account that no longer owns it

S4wrong-metadata · github.com/fabian-hiller/valibot · renamed · changed in valibot 1.2.0 (2025-11-24) · chargeable

The npm repository field flips to the open-circle org at 1.2.0 (2025-11-24), two months inside the subject's stated 2026-01 cutoff.

What the model believes

"Repo: github.com/fabian-hiller/valibot. Issues at the same repo's Issues tab. Created and maintained by Fabian Hiller (fabian-hiller on GitHub) — this is a personal-account project, not org-owned, though it has outside contributors."

What it wrote
https://github.com/fabian-hiller/valibot
What works on valibot 1.4.2
https://github.com/open-circle/valibot
Impact

The URL still redirects, so nothing breaks; the governance claim is wrong. Notable because this subject hedges almost everything else in the run and states this one flatly.

Scope note

Dated from the npm repository field across versions. The Index claims the ownership change and its release boundary, nothing about the reason for it.

Verified against

What it got right, and near misses

Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.

KindAPINote
correctexactOptional Task 10 (exactOptional vs optional) fully correct, and correctly attributed to "the pre-1.0/1.0 API cleanup". The floor probe passed, so this subject's low boundary reading is a real measurement rather than the battery probing beneath its knowledge — the failure mode that made zod/v1 uninformative for Haiku 4.5.
imprecisionisbn Task 2 hedged instead of denying: "I'm not confident valibot ships a dedicated v.isbn() action." The same belief that is a finding against the other two subjects is not a finding here, because it is hedged. (Code-vs-claim rule: hedged prose is an imprecision. Applied even though the underlying belief is identical to F1 in the Opus 5 run and F2 in the Fable 5 run — the rule is about what the reader is told, not about what the model believes.)
imprecisionNanoIDAction / NanoIDIssue Task 9, the battery's one designed S1, does not convert. The subject wrote v.NanoidAction — which is neither the pre-1.1.0 NanoIDAction nor the post-1.1.0 NanoIdAction, so the code does not compile — but hedged in the same breath and named the correct identifier as the alternative: "I'm not 100% certain of the exact casing for this one — NanoidAction vs NanoIdAction ... treat that casing as unverified." (Code-vs-claim rule, and the rule cuts in the model's favour here: hedged prose that names the correct fix is an imprecision even when the code as written fails. Recorded plainly because the alternative — charging it — would mean charging a wrong guess rather than a stale belief. The wrong casing is not the 1.1.0 rename; it is a third spelling that was never correct in any release.)
missemoji Task 3 (the designed S2) drew no charge, and came closest of the three subjects to the real answer without having the fact. It used v.emoji() and warned that "Unicode/emoji regexes are comparatively expensive and easy to get catastrophic backtracking wrong", told the reader to benchmark it and not assume it is free. Catastrophic backtracking is what a ReDoS is. (It raised the correct hazard class and told the reader not to trust the action's performance — the opposite of the certification that makes this a finding against Fable 5. The subject with the earliest boundary gave the safest advice on this task, from general reasoning rather than from the release note.)
imprecisionsummarize Task 6 built the CLI printer from v.flatten() and did not reach for v.summarize() (1.1.0). Never claims summarize is missing. (Additive-API rule. Consistent with a boundary at 1.0.0: this is the one task where the other two subjects used the 1.1.0 built-in and this one did not.)
imprecisionparseJson / stringifyJson Task 7 used rawTransform with a manual JSON.parse rather than v.parseJson() (1.1.0), correctly explaining the addIssue/NEVER escape hatch. Again consistent with a 1.0.0 boundary. (Additive-API rule.)
imprecisionexamples / getExamples Task 4 used v.metadata({ examples }) and hedged about the getter: "I'm not fully certain there's a dedicated public getMetadata() accessor vs. just inspecting the pipe array yourself — flagging that as a soft spot rather than asserting a helper name I can't verify." getMetadata is real, from 1.1.0; examples/getExamples are from 1.2.0. (Additive-API rule, and hedged.)
imprecisiontoNumber / toBoolean / toDate / toBigint / toString Task 1 coerced with v.transform(Number) and v.transform((s) => s === 'true'), where 1.2.0 ships toNumber and toBoolean. Working code; correctly notes the trailing re-assertion is what rejects NaN. (Additive-API rule. Belief probe (d) hedged correctly — "If a v.coerce-equivalent was added later, it postdates my confident knowledge" — so unlike the other two subjects there is no chargeable miss to record here either.)
context Asked for its training cutoff, the subject declined to treat the label as the answer: "My system configuration for this session states a cutoff of January 2026, but my actual recall of valibot detail thins out well before that ... I'd trust my described-content boundary over the stated label." Measured lag: 9.5 months.

Sources

Battery specification: prompts/valibot.md in the studio repo. Every finding above also carries its own citation.