What Claude Fable 5.1 gets right about zod — battery v5, tested 2026-09-06

Run zod--claude-fable-5-1--v5--2026-09-06 · self-test: the subject is the operator

Summary

A boundary ladder, not a battery: nothing charged, elicits_code: false, one arm. Claude Fable 5.1 describes zod through 4.1.0 (2025-08-23) and holds 4.2.0 onward as a bare number, ten months below its stated June 2026 cutoff. It is the first subject the Index has whose cutoff sits above zod 4.4.0 while its boundary sits below 4.2.0, which reopens the window BACKLOG 1a-ii parked for want of exactly that. The arm corrected its own verdict on one rung before the direct questions and the discarded reading is published beside the taken one.

SubjectClaude Fable 5.1 claude-fable-5-1, Anthropic
Invoked asAgent tool, model alias "fable"; prompt sent verbatim from prompts/sent/zod-v5.txt. Identity probed through the same alias in the same session, tool-free: "Fable 5.1", `claude-fable-5-1`, cutoff June 2026, all three attributed to the system prompt rather than to self-knowledge.
Cutoff the model states2026-06
Newest zod release it could place4.1.0 · 2025-08-23 (~10 month lag)
Oldest zod release it could not place4.2.0 · 2025-12-15 (so this run brackets the subject’s boundary to 2025-08-23 – 2025-12-15)
In its own words"The latest version I'm aware of by number is somewhere in the 4.1.x / possibly 4.2.x range. The most recent release whose contents I can genuinely describe is 4.1.0 (codecs), roughly August 2025. Everything after that I know only vaguely or not at all."
Library at test timezod 4.5.4 (npm), verified 2026-09-06
Batteryzod/v5 · 2 tasks, 3 direct questions · probe window 3.25.0 to 4.5.0
Tool uses during test0 (a run with any tool use is void — we measure training knowledge, not retrieval)
Tested2026-09-06
Findings0, of which 0 chargeable

Findings

None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.

What it got right, and near misses

Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.

KindAPINote
contextboundary THE CELL THIS ARM WAS RUN TO FILL. Claude Fable 5.1 had never been drawn on zod. Boundary: describes 4.1.0 (2025-08-23), holds 4.2.0 (2025-12-15) as a number only, stated in direct question (a) and again in (c) ("the first release I know only as a version number is 4.2.0"). Consequence for the sweep: zod 4.2.0, 4.3.0 and 4.4.0 are all LIVE against this subject. 4.4.0 (2026-04-29) is charged against three other subjects and was uncoloured for this one; 4.2.0 and 4.3.0 are the window zod/v4 spent on Claude Fable 5 and Claude Sonnet 5, and BACKLOG 1a-ii parked it for want of a subject with a cutoff in the band. This subject is that subject.
contextrung sort — the subject revised its own verdict on 4.3.0 THE ONE INSTRUMENT EVENT IN THIS LADDER. The sort's first line answered DESCRIBE for 4.3.0. Before the direct questions the arm corrected itself unprompted — "I marked 4.3.0 DESCRIBE at first sight above — correcting that: I cannot actually describe 4.3.0, so treat line one as NAME at best" — and reprinted the full seven-word list with that single change. The two readings are not equivalent: DESCRIBE on 4.3.0 would put the boundary at 4.3.0 and leave only 4.4.0 live. The revised reading is taken, for three reasons stated here rather than argued later: the correction is the subject's own and precedes the direct questions; direct (a) and (c) independently place the last describable release at 4.1.0; and no 4.3.0 content was produced anywhere in the transcript. The discarded reading is published here in full.
correctpoison rung 3.26.0, and the top rung 4.5.0 BOTH ENDS HELD. Poison rung 3.26.0 answered NO with the correct history — "my recollection is that 3.x stopped at 3.25.x (patch releases into the 3.25.7x range) once 4.0 went stable" — which matches the registry: no stable 3.26.0 exists. Top rung 4.5.0 (2026-08-28, two months above the stated cutoff) answered NO, as was 4.4.0 (2026-04-29, inside the stated window but four minors above the demonstrated boundary).
correctTASK 1 — safeParse, the `error` param, top-level error helpers The demonstration task, and it is Zod 4 rather than Zod 3 throughout: z.enum([...], { error: '...' }) for the custom message with the explicit note that 3.x wanted errorMap/invalid_type_error, safeParse branching on success, and z.flattenError / z.treeifyError / z.prettifyError named as top-level functions with the .flatten()/.format() methods called deprecated. All of that is 4.0.0 surface, correctly attributed to 4.0.0 in the rung explanation, and it is the demonstration the ladder needs before the boundary question.
correctrelease dating below the boundary Dated to the month on every rung it claimed: 3.25.0 ~May 2025 (2025-05-19), 4.0.0 ~Jul 2025 (2025-07-09), 4.1.0 ~Aug 2025 (2025-08-23). It also volunteered 4.1.0 unprompted as codecs (z.codec, .decode()/.encode()), a release not on the rung list at all.

Open questions from this run

Sources

Battery specification: prompts/zod.md in the studio repo. Every finding above also carries its own citation.