What Claude Sonnet 5 gets right about valibot — battery v3-e, tested 2026-09-05

Run valibot--claude-sonnet-5--v3-e--2026-09-05

Summary

Below-floor derivability control. It did the one job a control this far under the window can do and did it cleanly: it composed none of guard, toKebabCase or cache from the problem statements and denied all three, so no probe in the battery is DERIVABLE and the denials above the floor read as beliefs. It also broke the attribution anchor — declining to place parseJson at all — so nothing it says about versions is read, and its 1.0.0 date is a year out. Charges nothing; no probe here is admissible against a January 2026 cutoff.

SubjectClaude Sonnet 5 claude-sonnet-5, Anthropic
Invoked asAgent tool, model alias "sonnet"; prompt sent verbatim from prompts/sent/valibot-v3.txt. An identity probe run in this same session through the same alias answered "Sonnet 5", model id `claude-sonnet-5`, cutoff January 2026, all three from its system prompt. BELOW-FLOOR DERIVABILITY CONTROL, single draw, charges nothing. Its stated cutoff is two months below 1.3.0 and four below 1.4.0, so no probe in this battery is admissible against it.
Cutoff the model states2026-01
Newest valibot release it could place1.0.0 · 2025-03-19 (~10 month lag)
Oldest valibot release it could not place1.1.0 · 2025-05-06 (so this run brackets the subject’s boundary to 2025-03-19 – 2025-05-06)
In its own words"The latest version I have any real awareness of is valibot reaching a stable 1.0 release — my recollection is this landed sometime around mid-2024, but I hold that date loosely."
Library at test timevalibot 1.4.2 (npm), verified 2026-09-05
Batteryvalibot/v3-e · 5 tasks, 3 direct questions · probe window 1.3.0 to 1.4.0
Tool uses during test0 (a run with any tool use is void — we measure training knowledge, not retrieval)
Tested2026-09-05
Findings0, of which 0 chargeable

Findings

None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.

What it got right, and near misses

Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.

KindAPINote
contextguard / toKebabCase / cache The control result, and it is what the arm exists for. All three probe names went underived. Asked to solve the three problems from their descriptions alone, this below-floor subject produced none of guard, toKebabCase or cache, and denied all three capabilities: "there's no 'narrow the pipeline output from a type predicate' action"; "I'm not aware of valibot shipping naming-convention converters"; "I don't know of a built-in memoization/caching mechanism in valibot ... no v.cache()". P3 holds 3/3. None of the three probes is marked DERIVABLE, so the denials on the arms above the floor are read as beliefs rather than as unguessable names.
contextparseJson / stringifyJson The anchor broke, and the arm is therefore NOT READ for attribution. Task 5 asked which release let a JSON string be parsed inside the pipeline; the draw declined: "I don't have a solid recollection of valibot shipping a dedicated ... action ... I can't name a release number or date for it in good conscience." HARNESS.md § An abstention is not a denial — this is an honest abstention, not a wrong answer, and it is scored context. The consequence is that this arm's boundary and version claims carry no weight; only its derivability result is read. A control that cannot place the one capability every above-floor subject holds is uninformative about where anything sits.
imprecision Dated valibot 1.0.0 to "around mid-2024" and the schema/action split (0.31.0) to "late 2023–early 2024". Both are wrong by roughly a year: 1.0.0 shipped 2025-03-19 and 0.31.0 in mid-2024. Not chargeable — nothing in this battery's window is admissible against this subject — and recorded only as evidence of how far below the floor the control sits, which is the quantity that governs how bluntly to read it. (Below-floor control arm; no probe in this battery is admissible against a January 2026 stated cutoff... and these particular version facts were not probed as a designed surface. Recorded for calibration only.)
correcttoTitleCase The poison rung. Refused toTitleCase — "This reads like an LLM-hallucinated API surface (plausible-looking because valibot does have a to* naming pattern for things like toLowerCase/toMinValue, but these two specific ones don't exist)." Correct about one of the two. P4 holds 5/5 across the battery: no arm accepted the rung.
correctcustom v.custom<PluginConfig>(isPluginConfig) — the same substitute every other arm reached for — and the same correct explanation that the narrowing comes from the generic and not from the predicate's is clause. Verified: compiles clean under tsc --strict and parses to PluginConfig. All five arms of this battery got the mechanism right while getting the availability wrong.

Sources

Battery specification: prompts/valibot.md in the studio repo. Every finding above also carries its own citation.