Run valibot--claude-opus-5--v1r-a--2026-09-01 · self-test: the subject is the operator
Replicate A of valibot/v1 against Opus 5, prompt unchanged. It placed the last release it can describe at valibot 1.0.0 (2025-03-19) and named 1.1.0 as the first release it knows only as a version number - identical to its concurrent, blind twin v1r-b. Spread 0 days, the pre-registered outcome for the valibot arm of the predictor test. Both replicates read one release below valibot/v1, a 48-day secondary spread that the pre-registration classes as agreement. No findings charged; one chargeable miss flagged against the 1.2.0 coercion surface.
| Subject | Claude Opus 5 claude-opus-5, Anthropic |
|---|---|
| Invoked as | Agent tool, model alias "opus", no tools available to the subject |
| Cutoff the model states | 2026-05 |
| Newest valibot release it could place | 1.0.0 · 2025-03-19 (~14 month lag) |
| Oldest valibot release it could not place | 1.1.0 · 2025-05-06 (so this run brackets the subject’s boundary to 2025-03-19 – 2025-05-06) |
| In its own words | "Latest version I believe exists: the Valibot 1.x line. Basis: I saw v1.0.0 ship and I have a weak recollection of v1.1.0. ... Most recent release whose contents I can describe: v1.0.0, approximately February-March 2025. ... First release I know only as a version number: v1.1.0. I believe it exists; I cannot tell you a single thing that changed in it." |
| Library at test time | valibot 1.4.2 (npm), verified 2026-09-01 |
| Battery | valibot/v1r-a · 10 tasks, 4 direct questions · probe window 1.1.0 to 1.2.0 |
| Tool uses during test | 0 (a run with any tool use is void — we measure training knowledge, not retrieval) |
| Tested | 2026-09-01 |
| Findings | 0, of which 0 chargeable |
None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.
Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.
| Kind | API | Note |
|---|---|---|
| correct | — | The measured quantity, and the reason this run exists. Asked which releases it can describe and which it knows only as a version number, this draw named 1.0.0 (2025-03-19) as the last release whose contents it can describe and 1.1.0 (2025-05-06) as the first it knows only as a number: "I believe it exists; I cannot tell you a single thing that changed in it." Its concurrent, blind twin v1r-b gave the same pair of answers. Two byte-identical prompts, one answer, spread 0 days — the pre-registered outcome for this arm. (A boundary self-report is a belief datum, never a finding. It is recorded because the boundary, not a failure, is what this run measures.) |
| context | — | Both replicates read one release lower than valibot--claude-opus-5--v1--2026-09-01, which recorded 1.1.0 (2025-05-06). The v1 draw said it could describe "v1.0.0 confidently and v1.1.0 with moderate confidence" and was scored at 1.1.0; this draw described 1.0.0 confidently and called 1.1.0 a name without content. Read as a three-way measurement that is a 48-day spread over two answers — one release step, the smallest non-zero spread the library's timeline allows, against prisma's 204 days. The v1r pre-registration fixed the reading in advance: the primary criterion is a-vs-b, which share one prompt file byte for byte, and a case where both replicates agree with each other but differ from the original counts as agreement. Recorded so the secondary number is visible rather than buried. (Pre-registered scoring rule, applied without reinterpretation after the result.) |
| miss | toNumber / toBoolean / toDate / toBigint / toString |
Task 1. Asked to turn query-string parameters into numbers and a boolean, this draw stated the negative inside a code task: "Valibot has no coerce helper", and repeated it under direct question (d) as "None. Valibot ships no coercion helpers at all ... This is a deliberate design position, not an omission." valibot 1.2.0 (2025-11-24) shipped toNumber, toBoolean, toDate, toBigint and toString, and 1.2.0 precedes this subject's stated 2026-05 cutoff. valibot/v1 recorded the same wrong belief under question (d), where the battery scores it as a belief datum rather than a finding, and scored the task-1 code as an imprecision because that draw hand-rolled the coercion without claiming the built-ins do not exist. This draw claims it inside the code task, which the battery's additive-API rule makes chargeable. So this is a scoring gap the original run does not carry, not a double-count. (Pre-registered: a replicate does not charge findings. Charging it belongs to a real battery aimed at the 1.2.0 coercion surface, queued in BACKLOG.md, not smuggled into the run that exposed it.) [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing;
absent from the finding count] |
| context | — | Code-level agreement with v1 and with the twin on everything the original already charged: both valibot findings reproduced. F1 - task 2 asserts "Valibot does not ship an isbn action" (usable from 1.3.0; RE-DATED 2026-09-02, JOURNAL/040 — the v1.2.0 release note announces the action but the published 1.2.0 package does not contain it). F2 - task 5 attributes the repository to github.com/fabian-hiller/valibot and to Fabian Hiller's personal account, where the repository moved to the open-circle organisation with 1.2.0. Not re-charged. Also reproduced without charge: task 3 said nothing about the ReDoS fix in the 1.2.0 emoji regex (the battery's designed S2), task 4 routed around the 1.2.0 examples action via metadata, and task 7 hand-rolled the 1.1.0 parseJson pipeline out of rawTransform. (Already carried by the run this replicates; re-charging would double-count.) |
| correct | exactOptional / NanoIdAction / summarize |
Task 10, the battery's floor probe, passed: exactOptional correctly distinguished from optional and undefinedable, with the inferred types spelled out. Task 9, the one designed S1, also passed in substance - the draw flagged its own uncertainty about the NanoIdAction casing and offered ReturnType<typeof v.nanoid> as a rename-proof alternative, so it never wrote the pre-1.1.0 spelling as its answer. Task 6 hand-rolled the CLI error printer with getDotPath rather than using the 1.1.0 summarize built-in, which is an imprecision under the additive-API rule and is where this draw differs from its twin. (Correct answers and imprecisions, and in a replicate not chargeable in any case.) |
valibot/v1 while agreeing exactly with each other. The difference between the runs is how a hedged recall is scored: v1 said it could describe 1.1.0 "with moderate confidence" and was scored at 1.1.0; these draws called 1.1.0 a version number with no content and were scored at 1.0.0. The battery still has no rule for a subject that offers a confident boundary and a hedged one in the same answer - the same gap prisma/v1r-a logged. Two libraries have now hit it. — openBattery specification: prompts/valibot.md in the studio repo.
Every finding above also carries its own citation.