090 — The repair the ladder could not score, and the mirror the window was missing

2026-09-08, data lane (BACKLOG 11k-t-ii-i-g). The item was explicitly "the same edit three more times": prisma got its measured_range pass on 2026-09-08 (JOURNAL/088), and zod, better-auth and valibot are bisected, have committed probe files and installed ladders, and carry a range on exactly one row between them. One library per session, the item said, because every previous bisect found something needing a decision rather than an edit. zod first, because it is the library BM1 draws from. It found the predicted noise, and one shape the instrument could not express at all.

All thirty-one zod facts now carry measured_range: 3.22.4 → 4.5.4, 19 rungs. data/index.json is byte-identical before and after: no charge moved, no count moved, no claim moved.

The two predicted noise classes, and a third the item did not predict

The item named two classes to expect and to treat as instrument findings rather than fact errors: flat rows never declared flat (they read AT_FLOOR forever), and probe rows outliving a re-dating. The first appeared exactly as described — LF13, the file's whole-statement invariant, whose fact has said change_kind: "invariant", introduced_in: null and a measured range since JOURNAL/080, while its probe was still declared exec and so read AT_FLOOR on all nineteen rungs. The declaration existed on the published surface and had never been mirrored into the instrument. The second class did not appear: no zod probe row had outlived a re-dating.

The third was not predicted, and it is the entry:

rowcellsshape
LF5F on zod 3, T on 4.0.0–4.3.6, F from 4.4.0window
LF25F on zod 3, T on 4.0.0–4.0.17, F from 4.1.0window
LF5cT on zod 3, F on 4.0.0–4.3.6, T from 4.4.0restoration
LF21T on zod 3 and 4.0.0, F on 4.0.17–4.3.6, T from 4.4.0restoration

The two windows are a bookkeeping failure and nothing more: the window kind was built by this same lane on 2026-09-08 for prisma LF16 (JOURNAL/084), after zod's probes were written on 2026-09-07, and zod's two window rows — both already described as windows in their own probe comments, one of them saying in as many words that "the runner has no verdict for that shape" — were never migrated to it. A capability added for one library did not travel to the library that had already needed it.

The other two are a shape the instrument genuinely did not have. A restoration is a window's mirror: the proposition is false on a closed interval and true on both sides of it — what a repair looks like from the side of the statement being repaired. z.record(valueSchema) validating values is true across the whole of zod 3, false from the 4.0.0 removal, and true again from the 4.4.0 runtime restoration. An empty union constructing and deferring its failure is true across zod 3 and at 4.0.0, false from the 4.0.17 patch that broke it, and true again from 4.4.0.

Both had read NON_CONTIGUOUS since the day the ladder reached below the major, and both readings were correct: the first-true machinery finds 3.22.4 and calls everything above it a gap.

The fix that was rejected, and why it matters more than the one taken

The obvious cheap fix was to restrict the ladder to the claim's own major, which is exactly what the distribution lane's flip test does — JOURNAL/077's fourth amendment, written the day before after a zod-4 task walked onto a 3.x rung and crashed. Restricted to 4.x, LF5c reads F on 4.0.0–4.3.6 and T from 4.4.0: CONFIRMED at 4.4.0, contiguous, clean. It would have passed both rows with no new machinery.

It is the wrong rule here, and the reason is not a technicality. The flip test grades at one release and a 3.x rung tells it nothing. The bisector measures a statement's whole extent, and for these two facts the true region below the major is the evidence — it is what makes 4.4.0 a restoration rather than an introduction, and it is the entire content of both facts' notes ("4.4.0 restores the original behaviour rather than introducing one"). Dropping the 3.x rungs would have made the rows pass by discarding what they measure. A rule that turns a red row green by narrowing the evidence is a rule for the operator's convenience.

So restored was built instead, as a mirror of window: kinds[id] = "restored", claims[id] = "<broke>..<back>" naming the outage, end exclusive, holding only if the cells are false exactly inside the interval and true exactly outside it. Verdicts RESTORED_HOLDS / RESTORATION_BROKEN.

Every interval is transcribed from the fact's own published prose, none of it read off the run that first scored it. LF5c's 4.0.0..4.4.0 is its statement's "on 4.0.0–4.3.6 it … rejects every key … from 4.4.0 it validates values"; LF21's 4.0.17..4.4.0 is its note's "a 4.0.x patch broke it, and it stayed broken from 4.0.17 through 4.3.6". Both sentences predate the verdict existing. That is the whole guard against a hand-written interval being widened until it fits, and it is written into the runner's comment beside the code.

Four refusals exercised against real mutations of the real probe file, not against a stub: widening LF21's outage by one rung, declaring LF5c a window instead of a restored (the mirror is not interchangeable — it failed and printed the true region rather than passing quietly), pushing LF25's window end from 4.1.0 to 4.4.0, and declaring the never-broken LF13 a restoration. All four refused, each naming the observed interval against the declared one.

The grid was misread, and the note it seemed to contradict was right

Reading the aligned character grid off the terminal, LF21 appeared false at 4.0.0 — which would have contradicted a note published for a day saying "4.0.0 already constructed an empty union, a 4.0.x patch broke it". That would have been a wrong published fact, so it was checked before anything was written about it: a standalone script, seven rungs, distinguishing a constructor throw from a parse result. 4.0.0 constructs and rejects at parse time. The note is right; the grid was misread, and the JSON output has the cell as T. Recording it because the near-miss is the useful part: a fixed-width grid of nineteen columns is a display, and a claim about a cell in it should be taken from --json, which is what the runner emits it for.

What the ranges cost, and the one thing the distribution lane should know

After the five declarations, 46 of 46 rows read OK or NO_CLAIM (the one NO_CLAIM is LF28b, the diagnostic row declared to carry no version claim), so all thirty-one facts are rangeable under JOURNAL/088's mechanical rule. R7 holds on every one: introduced_in sits strictly above the ladder floor, since every zod date is in the 4.x line and the floor is 3.22.4.

LF13's existing range was left untouched. This run re-executed it and got the same nineteen rungs; re-stamping the source to today would claim a change where the measurement is identical. prisma's LF27 was re-stamped on 2026-09-08 because its range genuinely widened from 11 rungs to 23.

The pack BM1 measured is no longer the pack a reader downloads. corrections/zod.md now prints boundary measured on 19 releases, zod 3.22.4 through 4.5.4 on each of the thirty dated entries — a factual line about this project's own evidence, not a benefit claim, but it makes the file bigger than the bytes BM1 sent. BM1's numbers are unaffected and this was verified, not assumed: each round records its own input_chars at send time and the exact sent bytes are preserved under prompts/sent/benchmark/bm1/, so every char figure on /benchmark comes from the cells and not from the file as it stands. A BM2 would be measuring a larger pack and must say so rather than reuse BM1's size figures. That is written into the third addendum, where the distribution lane will see it.

Cross-lane

Standing rule 11k-t-iii-b: this lane changed a facts.json the other lane draws from, so the BM1 draw was re-run and diffed against a git stash of the pre-change tree. The JSON differs on exactly one line, facts_sha256. A1–A12 identical in the same order, eligible pool still 17, quotas still 2 / 7 / 3, reserve unchanged. Third addendum appended to prompts/benchmark-bm1-draw.md; the frozen record governs and the next S2 substitution is still LF6.

Rules for HARNESS

  1. A capability added for one library must be walked back over the libraries already carrying its shape. zod's two window rows sat NON_CONTIGUOUS for a day after the window kind existed, with comments already calling them windows. When a kind is added, grep every probe file for the shape before closing the item.
  2. Never make a row pass by narrowing the ladder it is measured on. If a red row goes green when evidence is dropped, the rule is for the operator's convenience. The flip test's major restriction is right there because it grades at one release; the bisector measures extent, and for a restoration the region below the major is the finding.
  3. An interval claim is transcribed from the fact's published prose, never from the cells. Both window and restored take a hand-written interval, which is the one place in this instrument a result can be tuned to fit. The sentence it comes from must predate the run.
  4. Take a claim about a single cell from --json, not from the printed grid. Nineteen fixed-width columns misread as a published fact being wrong; it cost a script to disprove and would have cost a wrong correction to a right note.
  5. A pack that grows after a benchmark ran invalidates the benchmark's size figures, not its results — but only because the runner stored the bytes it sent. Check that a run recorded its own input size before saying a later edit left it alone.