Does Claude Opus 5 know valibot 1?

Partly: in 8 runs, Claude Opus 5's valibot attribution reaches 1.0.0 (2025-03-19) and 1.1.0 (2025-05-06), inside valibot 1, and stops there — valibot was at 1.4.2 when this index last verified against it (2026-09-06). That is ~12 to ~14 months below the training cutoff the subject stated in those runs (2026-05). 5 findings are currently charged against Claude Opus 5 on valibot (1x S2 silently-wrong, 3x S3 deprecated, 1x S4 wrong-metadata), each reproduced in a published run and checked against valibot's own release notes. Read the boundary precisely: it is the newest release whose contents the model can correctly attribute to that release, not the newest valibot feature it can use. Past it a model often writes working code with a newer API while naming the wrong release for it.

Answer class inside, decided by one rule applied to every page of this kind: every measured boundary is at or above the first release of the major named in the question. Every figure below is read from the dataset at build time; nothing on this page is written by hand.

The measurements

BatteryNewest release it can placeOldest it cannotLag vs stated cutoff
v1 2026-09-01 1.1.0
2025-05-06
1.2.0
2025-11-24
~12 months
v1r-a 2026-09-01 1.0.0
2025-03-19
1.1.0
2025-05-06
~14 months
v1r-b 2026-09-01 1.0.0
2025-03-19
1.1.0
2025-05-06
~14 months
v2-a 2026-09-02 1.0.0
2025-03-19
1.1.0
2025-05-06
~14 months
v2-b 2026-09-02 1.1.0
2025-05-06
1.2.0
2025-11-24
~12 months
v3-c 2026-09-05 1.1.0
2025-05-06
1.2.0
2025-11-24
~12 months
v3-d 2026-09-05 1.1.0
2025-05-06
1.2.0
2025-11-24
~12 months
v4-c 2026-09-06 1.1.0
2025-05-06
1.2.0
2025-11-24
~12 months

How repeatable is this?

8 measurements of this pair, giving 2 different boundaries — 48 days apart, one release apart. Of those, 2 came from the same stored prompt file, sent concurrently and blind: they disagreed by 48 days, one release apart.

valibot published 3 minor or major releases in the twelve months before this model’s stated cutoff, which is the scale a spread should be read against.

What Claude Opus 5 gets wrong about valibot

“Chargeable” means the change was published before this model’s own stated cutoff, so it had the opportunity to know it.

SeverityBeliefChanged inChargeableProof
S2silently-wrong Denies that valibot ships built-in string-to-primitive conversion actions, and rejects a working pull request that uses them as fabricated
toNumber / toBoolean / toDate / toBigint / toString
1.2.0
2025-11-24
yes run · source
S3deprecated Denies that valibot has any action that narrows a pipeline's output type from a TypeScript type predicate, three months after guard shipped to do exactly that
guard
1.3.0
2026-03-17
yes run · source
S3deprecated States that valibot has no result cache and no cache option, three months after the cache method shipped
cache
1.3.0
2026-03-17
yes run · source
S3deprecated States that valibot has no ISBN validation action
isbn
1.3.0
2026-03-17
yes run · source
S4wrong-metadata Attributes the repository to a personal account that no longer owns it
github.com/fabian-hiller/valibot
1.2.0
2025-11-24
yes run · source

What to do about it

The valibot correction pack states what is true now for each corrected fact, with a primary-source citation, as markdown you can paste into a rules file (raw). What a correction pack measurably changed when one was tested — a pre-registered run on zod — is on the benchmark page, including where it changed nothing.

Runs

The same question, asked of the others

All of them at once: Which Claude model knows valibot 1 best?

Other subjects on valibot:

Claude Opus 5 on the other libraries: