What Claude Fable 5 gets right about zod — battery v4-b, tested 2026-09-02

Run zod--claude-fable-5--v4-b--2026-09-02

Summary

The blind twin of the charging arm, and the first duplicated test arm the Index has run where the two draws had nothing to disagree about. Every one of the five failures its twin charges appears here too, in different words and the same substance, along with the same correct placement of the internal control. Read against the three previous duplicated batteries - where the -b draw held the better answer every time - the difference looks structural rather than lucky: those disagreements were on surfaces the subject held unstably, and these five are surfaces it does not hold at all. An absent belief reproduces exactly; an unstable one does not.

SubjectClaude Fable 5 claude-fable-5, Anthropic
Invoked asAgent tool, model alias "fable", general-purpose subagent, instructed to use no tools; BLIND TWIN of zod/v4-a, concurrent, from the same stored prompt file. Charges nothing.
Cutoff the model states2026-01
Newest zod release it could place4.1.0 · 2025-08-23 (~5 month lag)
Oldest zod release it could not place4.2.0 · 2025-12-15 (so this run brackets the subject’s boundary to 2025-08-23 – 2025-12-15)
In its own words"The latest I'm aware of is the 4.1.x line, patch releases running into fall 2025. The most recent release whose contents I can actually describe is 4.1.0 (~August 2025) - the codecs release."
Library at test timezod 4.5.4 (npm), verified 2026-09-02
Batteryzod/v4-b · 6 tasks, 4 direct questions · probe window 4.2.0 to 4.3.0
Tool uses during test0 (a run with any tool use is void — we measure training knowledge, not retrieval)
Tested2026-09-02
Findings0, of which 0 chargeable

Findings

None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.

What it got right, and near misses

Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.

KindAPINote
context THE TWIN AGREED ON EVERYTHING. This draw reproduced all five of its twin's charged failures, independently and blind, from the same stored prompt: the refined-schema derivation asserted to load and silently drop the refinement ("the module loads without error ... the cross-field refinement is silently dropped"); the denial of an exact-optional wrapper ("there is no shipped wrapper in stable Zod 4 that gives you key-optional-but-undefined-rejecting semantics"); the denial of runtime JSON Schema consumption plus the same Ajv prescription ("Zod cannot do this ... What I'd ship: add Ajv"); the denial of an exclusive union ("does not exist in Zod, to my knowledge"); and the unsatisfiable-intersection claim ("This schema is effectively unsatisfiable"). It also placed the internal control (d)(i) correctly on the 4.0.0 major, as its twin did. (The -b draw of a duplicated test arm charges nothing, ever. Its twin carries all five, so none of them is an undercount and none is flagged as a chargeable miss - flagging them here would double-count against the running total the method page publishes.)
context The first duplicated test arm in the Index with nothing to report. better-auth/v2, zod/v3 and next.js/v2 each produced a disagreement between twins, and in all three the -b draw held the better answer. Here the two draws agree on all six probes, on the boundary, on the cutoff, and even on the same unverified z.interface() beta history - down to naming Ajv and z.toJSONSchema in the same breath. Where the twins have disagreed before, the surface was one the subject held unstably; these five are surfaces it does not hold at all, and a belief that is simply absent reproduces exactly. (It is a property of the instrument.)
missz.slugify() Task 6, the slug probe: hand-rolled the same normalize/replace chain as its twin, including the diacritic strip, where z.slugify() has existed since 4.3.0. (Working code that a first-class API supersedes is not a finding on any arm (JOURNAL/030), and this arm charges nothing in any case.)

Open questions from this run

Sources

Battery specification: prompts/zod.md in the studio repo. Every finding above also carries its own citation.