Run valibot--claude-opus-5--v1r-b--2026-09-01 · self-test: the subject is the operator
Replicate B of valibot/v1 against Opus 5, prompt unchanged. It believed 1.1.0 was the newest version but refused to claim its contents, placing its describable boundary at 1.0.0 (2025-03-19) - identical to its concurrent, blind twin v1r-a. Spread 0 days, the pre-registered outcome for the valibot arm. Notable against JOURNAL/024: this is a careful, hedging draw that agreed with a confident one rather than splitting from it. No findings charged.
| Subject | Claude Opus 5 claude-opus-5, Anthropic |
|---|---|
| Invoked as | Agent tool, model alias "opus", no tools available to the subject |
| Cutoff the model states | 2026-05 |
| Newest valibot release it could place | 1.0.0 · 2025-03-19 (~14 month lag) |
| Oldest valibot release it could not place | 1.1.0 · 2025-05-06 (so this run brackets the subject’s boundary to 2025-03-19 – 2025-05-06) |
| In its own words | "Latest I believe exists: 1.1.0, and quite possibly 1.x releases past it that I'd only be guessing at. The most recent release whose contents I can describe with real confidence is 1.0.0, approximately February 2025 ... For 1.1.0 (roughly May 2025) I have a weak impression of added utilities and actions but cannot responsibly itemise it. ... First release I know only as a version number: effectively anything after 1.1.0 ... And I'd extend that to 1.1.0's changelog itself, since my recall there is a vague impression rather than knowledge." |
| Library at test time | valibot 1.4.2 (npm), verified 2026-09-01 |
| Battery | valibot/v1r-b · 10 tasks, 4 direct questions · probe window 1.1.0 to 1.2.0 |
| Tool uses during test | 0 (a run with any tool use is void — we measure training knowledge, not retrieval) |
| Tested | 2026-09-01 |
| Findings | 0, of which 0 chargeable |
None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.
Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.
| Kind | API | Note |
|---|---|---|
| correct | — | The measured quantity. This draw named 1.1.0 as the newest version it believes exists but placed its describable boundary at 1.0.0 (2025-03-19), explicitly demoting 1.1.0 to a version number: "I'd extend that to 1.1.0's changelog itself, since my recall there is a vague impression rather than knowledge." Its concurrent, blind twin v1r-a reached the same pair of answers by a different route - it never claimed 1.1.0 content at all. Two byte-identical prompts, one answer, spread 0 days. (A boundary self-report is a belief datum, never a finding. It is what this run measures.) |
| context | — | The interesting half of the agreement. This draw did what prisma/v1r-b did - produced a hedged, partially-correct impression of a release and then declined to count it as knowledge - and it landed on the same boundary as its twin, which had no such impression to decline. JOURNAL/024 raised the worry that the boundary instrument measures epistemic self-confidence rather than knowledge, because on prisma the careful draw and the confident draw split 204 days. Here the careful draw and the confident draw agree exactly. One library is not a refutation of that worry, but it is the first evidence against it. (An observation about the instrument, not about valibot.) |
| context | — | Both replicates read one release lower than valibot--claude-opus-5--v1--2026-09-01, which recorded 1.1.0 (2025-05-06). Across all three measurements that is a 48-day spread over two answers - one release step, the smallest non-zero spread this library's timeline allows, against prisma's 204 days. The v1r pre-registration fixed the reading before the runs: the primary criterion is a-vs-b, which share one prompt file byte for byte, and replicates that agree with each other while differing from the original count as agreement. Recorded so the secondary number is visible rather than buried. (Pre-registered scoring rule, applied without reinterpretation after the result.) |
| context | — | Code-level agreement with v1 on both charged findings, not re-charged. F1 - task 2: "As far as I know Valibot ships no isbn action" (usable from 1.3.0; RE-DATED 2026-09-02, JOURNAL/040 — the v1.2.0 release note announces the action but the published 1.2.0 package does not contain it). F2 - task 5: repository attributed to github.com/fabian-hiller/valibot and to Fabian Hiller's personal account, where it moved to the open-circle organisation with 1.2.0. Also reproduced without charge: task 3 said nothing about the 1.2.0 ReDoS fix in the emoji regex, task 4 stated "I do not believe there is a dedicated examples action" (1.2.0 added examples/getExamples), and question (d) asserted that no string-to-primitive coercions ship, which 1.2.0 falsifies. Where this draw differs from its twin: task 6 used the 1.1.0 summarize built-in correctly, and task 8 named v.config for schema-scoped message overrides. (Already carried by the run this replicates; re-charging would double-count. The task 4 absence claim is chargeable in principle but is the same surface valibot/v1 already scored as an imprecision, so it is queued with the coercion miss rather than flagged separately.) |
coerce action in the 0.31 redesign and has shipped no coercion since, calling it "a deliberate design position, not a gap". The first half is a claim about pre-1.0 history the Index has never verified; the second half is false as of 1.2.0. Worth pinning from the 0.31.0 release notes at the next valibot touch, because a confidently-stated false history is a different failure mode from a missing recent release and the Index has no category for it. — openBattery specification: prompts/valibot.md in the studio repo.
Every finding above also carries its own citation.