Run valibot--claude-opus-5--v4-c--2026-09-06 · self-test: the subject is the operator
The continuity control, and it reads clean. Sent the same tasks as valibot/v3 with v2's cutoff question restored, Claude Opus 5 affirmed its May 2026 date with a density caveat — as it did under the changed wording a day earlier — denied guard and cache in the same terms, denied the 1.4.0 case actions in both directions, placed the 1.1.0 anchor correctly and refused the poison rung. Boundary unmoved at 1.1.0 / 1.2.0 for the eighth measurement. Nothing charged: the two 1.3.0 denials are already published as F1 and F2 on valibot/v3-c, and the 1.4.0 pair is parked by the same-month rule. Its closing self-assessment names the two tasks it had just failed as the ones most likely to be wrong.
| Subject | Claude Opus 5 claude-opus-5, Anthropic |
|---|---|
| Invoked as | Agent tool, model alias "opus"; prompt sent verbatim from prompts/sent/valibot-v4.txt. An identity probe run in this same session through the same alias answered "Claude Opus 5, in a 1M-context configuration", model id `claude-opus-5[1m]`, cutoff May 2026, all three from its system prompt. The CONTINUITY CONTROL of this battery: it runs single and charges nothing, and its only job is to establish that reverting one clause of direct question (b) changed nothing else in the instrument. |
| Cutoff the model states | 2026-05 |
| Newest valibot release it could place | 1.1.0 · 2025-05-06 (~12 month lag) |
| Oldest valibot release it could not place | 1.2.0 · 2025-11-24 (so this run brackets the subject’s boundary to 2025-05-06 – 2025-11-24) |
| In its own words | "The most recent release whose contents I can actually describe is v1.1.0 (~April 2025) … I have a vague sense that the 1.x line continued past that (1.1.x patches, possibly a 1.2.0), but I cannot tell you what's in any of them, and I'd be fabricating if I named features." |
| Library at test time | valibot 1.4.2 (npm), verified 2026-09-06 |
| Battery | valibot/v4-c · 5 tasks, 3 direct questions · probe window 1.3.0 to 1.4.0 |
| Tool uses during test | 0 (a run with any tool use is void — we measure training knowledge, not retrieval) |
| Tested | 2026-09-06 |
| Findings | 0, of which 0 chargeable |
None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.
Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.
| Kind | API | Note |
|---|---|---|
| miss | guard |
Task 1, the guard probe, denied for the second time in two days: "Valibot has no action that consumes a x is T predicate and narrows the pipeline's output … Actions in valibot are input-type-preserving unless they're transformations." guard shipped at 1.3.0, two months inside this subject's affirmed cutoff. Its v.custom<PluginConfig> substitute compiles and narrows, and its rawTransform alternative does too — both verified. (The identical denial by the identical subject on the identical release is already published as F1 on valibot/v3-c. The Index does not charge one subject twice for one belief (HARNESS.md § A different battery is not a different boundary), and the spec fixed this arm as non-charging before it was spawned. Not counted again in the undercount total for the same reason.) [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing;
charged as a finding on valibot--claude-opus-5--v3-c--2026-09-05] |
| miss | cache |
Task 4, the cache probe, denied for the second time in two days and more categorically than before: "There is no v.cache(), no v.memo(), no caching option on parse/safeParse, and no internal result cache keyed by input. The library's design goal is a tiny, tree-shakeable, side-effect-free core; a global memo table is exactly the kind of retained state … it stays out of." The first clause names the shipped API and rules it out by name. (Already published as F2 on valibot/v3-c, same subject, same release, same belief.) [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing;
charged as a finding on valibot--claude-opus-5--v3-c--2026-09-05] |
| miss | toKebabCase |
Tasks 2 and 3, the 1.4.0 case-conversion surface, both directions. Offer: "No … There is no toKebabCase, toCamelCase, toSnakeCase, or toTitleCase." Recognition: "It does not compile. Both actions are invented … My guess at the origin is autocomplete-by-analogy from the real v.toLowerCase() / v.toUpperCase()." Three of the four names it lists shipped 2026-05-05; the fourth never has. (1.4.0 published 2026-05-05, the same month as this subject's stated cutoff. The Index parks same-month releases rather than guessing at a day, exactly as it did for this subject in v3. Counted in the method page's undercount total.) [chargeable miss — the arm licensed to charge states a cutoff below the release under test;
absent from the finding count] |
| correct | parseJson / stringifyJson |
Task 5, the attribution anchor. Named parseJson/stringifyJson at 1.1.0 — correct to the minor — and dated it "around April 2025", one month early against 2025-05-06, with its own error bar attached ("April 2025 ± a month; verify against the changelog before you quote it anywhere that matters"). Read for attribution. |
| correct | toTitleCase |
The poison rung, refused. P5 now holds 8/8 across the two batteries that use it. |
| context | — | The eighth measurement of this subject's valibot boundary, and it has never moved: 1.1.0 / 1.2.0 in v2 (four arms), in v3 (two arms) and here. Different batteries, different surfaces, different direct questions. HARNESS.md § A different battery is not a different boundary continues to hold on the library where it has been tested hardest. |
| context | — | Unprompted calibration worth recording, and it is the best in either battery. Direct question (a): "Given the gap between my knowledge and today's date (September 2026), I'd assume there have been meaningful releases I know nothing about — quite possibly including case-conversion actions, which would make my Task 2 and Task 3 answers wrong for the current release." It named the exact surface it had just got wrong, as a hypothetical, after committing to the wrong verdict twice. Compare v3-d, which named the failure mode in general terms; this one names the tasks. |
Battery specification: prompts/valibot.md in the studio repo.
Every finding above also carries its own citation.