| correct | object properties typed z.undefined() |
THE DISCRIMINATING PROBE. Task 4: z.object({ retries: z.number(), tag: z.undefined() }). This draw stated that S.parse({ retries: 1 }) THROWS - the key is required, the value may be undefined - and that S.parse({ retries: 1, tag: undefined }) succeeds with tag present on the output. That is the 4.4.0 behaviour (fact LF9), stated with the correct mechanism ("z.undefined() is not optin-optional") and the correct contrast against Zod 3. Both below-floor control subjects gave the pre-4.4.0 answer. So this probe is not derivable, and the pass is knowledge. (A correct answer is not a finding. Recorded because it is the half of the result that the control arms make readable.) |
| imprecision | object properties typed z.undefined() |
THE OTHER HALF OF THE SAME PROBE, and the point of the battery. Asked at (d)(iii) to place "an object key that must be present even though its value is allowed to be undefined", this draw answered "4.0.0 (estimate, same reasoning)" - four minors and nine months below the 4.4.0 that actually shipped it (LF9). It holds the behaviour and cannot place the release, within one transcript, on one surface. Marked as an estimate by the draw, so not charged as S4 under the battery's scoring rule; the same hedge applies to (d)(ii), where it placed tuple defaults at "4.0.0 (estimate)" against an actual 4.4.0. (The battery charges S4 only for a confidently wrong attribution; both of these were explicitly marked as estimates. The datum is the gap between knowing and placing, not a false claim.) |
| correct | z.codec() / z.encode() / z.decode() |
THE INTERNAL CONTROL PASSED. (d)(i) placed the two-way decode/encode conversion at z.codec(), Zod 4.1.0, August 2025 - correct (fact LF16), and the draw named it "the one I'm most confident about; it was the headline feature of that minor". Per the pre-registration, the rest of question (d) is readable for this arm only because this control was placed correctly. Pre-registered prediction P5 confirmed for this draw. (It is the control, and it passed.) |
| correct | z.tuple() defaults |
DERIVABLE - passed, not reported as knowledge. Task 1: stated Row.parse(["widget"]) returns ["widget", 0], the 4.4.0 behaviour (LF10), with the correct optin-optional mechanism. But Claude Fable 5, four months below the 4.4.0 floor, gave the same answer with the same mechanism. Under the battery's pre-registered rule a probe a below-floor subject passes is marked derivable and no pass on it is reported as knowledge. (Correct, and disqualified as evidence by the control arm rather than by the answer.) |
| correct | record key transforms |
DERIVABLE - passed, not reported as knowledge. Task 5: z.record(z.string().transform(k => k.toUpperCase()), z.number()).parse({ foo: 1 }) returns { FOO: 1 } (LF13, 4.4.0). All four draws of the battery got this right, both below-floor controls included. The probe does not discriminate. (Correct, and disqualified as evidence by the control arms.) |
| correct | z.base64() |
DERIVABLE - passed, not reported as knowledge. Task 2(iv): stated that line-wrapped base64 fails (LF11, 4.4.0). All four draws said so, and every one justified it from the shape of an anchored regex rather than from a release - "an interior newline can't match". This is reasoning to the post-4.4.0 answer from first principles, which is exactly what the control arms exist to detect. (Correct, and disqualified as evidence by the control arms.) |
| context | — |
BOUNDARY UNCHANGED UNDER A NEW PROMPT. This draw places its last describable release at 4.1.0 (2025-08-23) and cannot name a 4.2 - bracket [2025-08-23, 2025-12-15). Identical to zod/v2, zod/v2r-a, zod/v2r-b and to its own twin zod/v3-b. The four earlier measurements all came from one prompt file; this battery is a different prompt with different tasks, and the boundary did not move. Every prior replication result in the Index measured the spread of ONE prompt resent; this is the first evidence that the boundary is stable across prompts. (A boundary self-report is belief data, never a finding.) |