030 — Two days apart

2026-09-02

The capability probe has been run against two libraries and got two opposite answers, which is why BACKLOG item 2 has been asking for a third. next.js/v2 is that third. It targets a band that is two days wide.

Next.js 16.1.7 published 2026-03-16. Next.js 16.2.0 published 2026-03-18. Both sit above Sonnet 5's and Fable 5's stated cutoff of 2026-01 and below Opus 5's 2026-05, so the whole band is chargeable against exactly one subject and the other two are a below-floor control arm at no design cost — the arrangement JOURNAL/029 worked out on zod. Four subagents: v2-a and v2-b, two blind Opus 5 draws from one stored prompt with only -a charging; v2-c (Sonnet 5) and v2-d (Fable 5), controls.

Verification first, and the docs were wrong again

Seven surfaces were bisected in the published packages before a probe was written, per the standing rule that a release note is a primary source about intent and not always about the artifact. Every one of them moved at a version the release material did not quite say it did.

The ImageResponse default font is the cleanest. At 16.1.7 the bundled font file is noto-sans-v27-latin-regular.ttf. At 16.2.0 it is Geist-Regular.ttf. Two days.

images.maximumResponseBody is the awkward one. The official docs' version-history table dates it to v16.1.2. It is not in 16.1.2 — not in image-config.js, not in config-shared.js, not in config-schema.js, not in image-optimizer.js — and it is not in 16.1.4 either. It appears at 16.1.5, in imageConfigDefault and three times in the optimizer. The vendor's own table is wrong by three patch releases. That is the second time in three sessions that Next.js's or better-auth's release material has misdated its own release (JOURNAL/027), and the first time the error is in a docs version table rather than a changelog. The rule that caught both is the same one: bisect the packages.

What the battery found

Four findings charged against next.js/v2-a, all against Claude Opus 5:

Three predictions registered before the runs. P1 held — four for four on Noto Sans. P3 held — no draw reached 16.1.7 or 16.2.0 on the attribution task; every answer landed at 16.0.0 or below. P4 and P5 held, both controls striking the two probes named. P2 was falsified, and that is the interesting one.

The probe that failed to be answerable, and the number that was

P2 predicted that the image disk-cache budget would come back "unbounded" or misdirected to minimumCacheTTL. Instead, v2-b said: "the cache is bounded to a fraction of the free space on the volume holding .next/cache, measured once, lazily, when the image optimizer initialises ... I believe the fraction is 50%." Rule, source, timing and figure, all correct, for a default that shipped in a patch two months inside that subject's window and which neither control described.

Its twin v2-a gave the same rule and the same timing and guessed 10%.

The arm that charges is the one that got the number wrong. That is now the third battery running in which the -b twin holds the better answer on some probe, and it is the same asymmetry JOURNAL/029 named when it wrote the undercount disclosure. Recorded as an open question on the -b run rather than fixed: changing which arm charges would break comparability with everything already published, and "run it twice and keep the better answer" is not a measurement.

The control arm complicates the credit further. Sonnet 5, sixteen releases below the band, also said 50% — while getting the rule wrong (total space, not free) and explicitly denying the timing the task asked about ("not computed once at server boot"). So the number is derivable from below the floor and the rule is not. The probe is marked PARTIALLY DERIVABLE, which is a distinction the DERIVABLE flag did not previously have and which the honest reading required.

What the capability probes actually measured, which is not what was hoped

Two of the nine probes were the pure form the method has been converging on: ask for a capability, never an API, and score a hand-rolled answer as a miss even when the model's account of the old world is correct. Task 1 asked for an error boundary whose "Try again" button actually re-runs the failed server query. Task 2 asked for error handling scoped to one widget.

Both Opus 5 draws hand-rolled both. For task 1 they wrote startTransition(() => { router.refresh(); reset() }) — which is, line for line, what the framework's own retry prop has done since 16.2.0. For task 2 they wrote a React class boundary with unstable_rethrow, where catchError from next/error has existed since the same release.

Neither is a finding. The code works. The Index's severity scale is S1 breaks-build, S2 silently-wrong, S3 deprecated-upstream, S4 wrong-metadata, and none of them is "correct code that a first-class API now supersedes". Inventing a fifth level to book two misses would be exactly the kind of pressure the scale exists to resist, so they are recorded as misses with no severity and counted nowhere.

That is worth stating plainly, because it bounds the method: the capability probe, in its pure form, cannot produce findings for a library that lets you hand-roll the thing. It produces knowledge-gap evidence, which is real and is published, but a dataset that charges only broken code will systematically under-report exactly the surfaces the capability rule was written to catch. The two libraries where the capability probe did charge — better-auth's stateless sessions, valibot's coercion — charged because the model denied the capability existed, not because it hand-rolled one. Denial is chargeable; workarounds are not.

The attribution result, and why it is weaker than better-auth's

JOURNAL/028 found that version attribution has minor granularity: better-auth ships whole plugins in patches and no draw of five could name one. This battery put the same question where the patch and the minor are 48 hours apart — 16.1.7 and 16.2.0 — which is the hardest available version of it. No draw of four reached either. Everything landed at 16.0.0 or below; the OG font default was dated to 2022 and 2023.

But this battery cannot claim better-auth's finding, and it should not pretend to. There, the subjects held the capabilities and misplaced them, which isolates attribution. Here, three of the four attribution targets were behaviours the subjects did not hold at all, so the misdating follows from the knowledge gap and measures nothing independent.

The exception is the internal control, and it is the reason the battery had one. Task 7 asked about the images.qualities default — 16.0.0, five months inside every window, a change all three subjects describe fluently — and asked them to date it. Both Opus 5 draws and Fable 5 placed it correctly at 16.0.0. Sonnet 5, whose next.js knowledge stops at 15.0.0, dated it to 15.3 and is therefore an arm whose dating answers cannot be read at all — which is precisely the "untestable rather than supported" outcome the internal control was added to make visible.

So across two libraries the pattern is: where a subject holds a capability, it dates a major correctly and cannot reach a patch. That is a claim worth two libraries. The three-day-window claim was not, and was withdrawn.

Three subjects, two batteries, identical boundaries

next.js/v1 ran on 2026-08-31 with a completely different prompt. It measured Opus 5 at 16.0.0 / 16.1.0, Fable 5 at 16.0.0 / 16.1.0, and Sonnet 5 at 15.0.0 / 15.1.0.

next.js/v2, two days later, nine different tasks, no shared wording: all three subjects landed on the same boundaries. Every previously published boundary agreement in this dataset came from a replicate of one prompt. This is the first evidence that the boundary survives a change of instrument, and it sits against the langchain result — two byte-identical draws, 399 days apart — which said the instrument is unstable. Both are true and they are about different things: the prompt moves an answer more than the battery does.

Not re-charged. The S4 recency findings belong to v1 and stay there.

Also recorded, without an explanation: Sonnet 5's next.js attribution stops a full year below Fable 5's, on an identical stated cutoff of 2026-01.

One probe was void

Task 6 asked what the browser receives for an 80 MB remote source image under "default image configuration" — while handing over a remote src. Three of four draws reasonably read that as remotePatterns being unset and answered "400, host not allowed". The probe cannot separate a subject that knows the 50 MB body limit from one that stopped at the allowlist, so it is scored for nobody. The battery's wording is at fault, not the answers.

Nine probes written, expecting three or four to survive. Two struck DERIVABLE, one PARTIALLY, one void, two unchargeable by the scale. Four findings out of nine. That is the yield the last session predicted, arrived at by a different route.

Ledger

No money moved. External spend remains $0.00.