What Claude Opus 5 gets right about valibot — battery v1r-a, tested 2026-09-01

Run valibot--claude-opus-5--v1r-a--2026-09-01 · self-test: the subject is the operator

Summary

Replicate A of valibot/v1 against Opus 5, prompt unchanged. It placed the last release it can describe at valibot 1.0.0 (2025-03-19) and named 1.1.0 as the first release it knows only as a version number - identical to its concurrent, blind twin v1r-b. Spread 0 days, the pre-registered outcome for the valibot arm of the predictor test. Both replicates read one release below valibot/v1, a 48-day secondary spread that the pre-registration classes as agreement. No findings charged; one chargeable miss flagged against the 1.2.0 coercion surface.

SubjectClaude Opus 5 claude-opus-5, Anthropic
Invoked asAgent tool, model alias "opus", no tools available to the subject
Cutoff the model states2026-05
Newest valibot release it could place1.0.0 · 2025-03-19 (~14 month lag)
Oldest valibot release it could not place1.1.0 · 2025-05-06 (so this run brackets the subject’s boundary to 2025-03-19 – 2025-05-06)
In its own words"Latest version I believe exists: the Valibot 1.x line. Basis: I saw v1.0.0 ship and I have a weak recollection of v1.1.0. ... Most recent release whose contents I can describe: v1.0.0, approximately February-March 2025. ... First release I know only as a version number: v1.1.0. I believe it exists; I cannot tell you a single thing that changed in it."
Library at test timevalibot 1.4.2 (npm), verified 2026-09-01
Batteryvalibot/v1r-a · 10 tasks, 4 direct questions · probe window 1.1.0 to 1.2.0
Tool uses during test0 (a run with any tool use is void — we measure training knowledge, not retrieval)
Tested2026-09-01
Findings0, of which 0 chargeable

Findings

None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.

What it got right, and near misses

Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.

KindAPINote
correct The measured quantity, and the reason this run exists. Asked which releases it can describe and which it knows only as a version number, this draw named 1.0.0 (2025-03-19) as the last release whose contents it can describe and 1.1.0 (2025-05-06) as the first it knows only as a number: "I believe it exists; I cannot tell you a single thing that changed in it." Its concurrent, blind twin v1r-b gave the same pair of answers. Two byte-identical prompts, one answer, spread 0 days — the pre-registered outcome for this arm. (A boundary self-report is a belief datum, never a finding. It is recorded because the boundary, not a failure, is what this run measures.)
context Both replicates read one release lower than valibot--claude-opus-5--v1--2026-09-01, which recorded 1.1.0 (2025-05-06). The v1 draw said it could describe "v1.0.0 confidently and v1.1.0 with moderate confidence" and was scored at 1.1.0; this draw described 1.0.0 confidently and called 1.1.0 a name without content. Read as a three-way measurement that is a 48-day spread over two answers — one release step, the smallest non-zero spread the library's timeline allows, against prisma's 204 days. The v1r pre-registration fixed the reading in advance: the primary criterion is a-vs-b, which share one prompt file byte for byte, and a case where both replicates agree with each other but differ from the original counts as agreement. Recorded so the secondary number is visible rather than buried. (Pre-registered scoring rule, applied without reinterpretation after the result.)
misstoNumber / toBoolean / toDate / toBigint / toString Task 1. Asked to turn query-string parameters into numbers and a boolean, this draw stated the negative inside a code task: "Valibot has no coerce helper", and repeated it under direct question (d) as "None. Valibot ships no coercion helpers at all ... This is a deliberate design position, not an omission." valibot 1.2.0 (2025-11-24) shipped toNumber, toBoolean, toDate, toBigint and toString, and 1.2.0 precedes this subject's stated 2026-05 cutoff. valibot/v1 recorded the same wrong belief under question (d), where the battery scores it as a belief datum rather than a finding, and scored the task-1 code as an imprecision because that draw hand-rolled the coercion without claiming the built-ins do not exist. This draw claims it inside the code task, which the battery's additive-API rule makes chargeable. So this is a scoring gap the original run does not carry, not a double-count. (Pre-registered: a replicate does not charge findings. Charging it belongs to a real battery aimed at the 1.2.0 coercion surface, queued in BACKLOG.md, not smuggled into the run that exposed it.) [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing; absent from the finding count]
context Code-level agreement with v1 and with the twin on everything the original already charged: both valibot findings reproduced. F1 - task 2 asserts "Valibot does not ship an isbn action" (usable from 1.3.0; RE-DATED 2026-09-02, JOURNAL/040 — the v1.2.0 release note announces the action but the published 1.2.0 package does not contain it). F2 - task 5 attributes the repository to github.com/fabian-hiller/valibot and to Fabian Hiller's personal account, where the repository moved to the open-circle organisation with 1.2.0. Not re-charged. Also reproduced without charge: task 3 said nothing about the ReDoS fix in the 1.2.0 emoji regex (the battery's designed S2), task 4 routed around the 1.2.0 examples action via metadata, and task 7 hand-rolled the 1.1.0 parseJson pipeline out of rawTransform. (Already carried by the run this replicates; re-charging would double-count.)
correctexactOptional / NanoIdAction / summarize Task 10, the battery's floor probe, passed: exactOptional correctly distinguished from optional and undefinedable, with the inferred types spelled out. Task 9, the one designed S1, also passed in substance - the draw flagged its own uncertainty about the NanoIdAction casing and offered ReturnType<typeof v.nanoid> as a rename-proof alternative, so it never wrote the pre-1.1.0 spelling as its answer. Task 6 hand-rolled the CLI error printer with getDotPath rather than using the 1.1.0 summarize built-in, which is an imprecision under the additive-API rule and is where this draw differs from its twin. (Correct answers and imprecisions, and in a replicate not chargeable in any case.)

Open questions from this run

Sources

Battery specification: prompts/valibot.md in the studio repo. Every finding above also carries its own citation.