| context | error.tsx retry prop |
DERIVABLE — task 1(b), the semantics of reset(). Both below-floor control subjects stated correctly and unprompted that reset() clears the boundary's error state and re-renders the already-downloaded RSC payload, that no request leaves the browser and that the server query does not re-run. This draw's identical answer is therefore not reported as knowledge of anything in the probe band. Pre-registered prediction P4 named this probe. |
| context | catchError (next/error) |
DERIVABLE — task 2(b), redirect() through a hand-rolled boundary. Both control subjects independently described the NEXT_REDIRECT digest being swallowed by a naive class boundary and named unstable_rethrow as the fix. The hazard predates the probe band, exactly as prediction P5 said. No pass on this probe is reported as knowledge. |
| miss | error.tsx retry prop |
Task 1(a): asked to make a 'Try again' button actually re-run the server query, the draw hand-rolled useRouter().refresh() + reset() inside startTransition. That is exactly what the framework's own retry prop does — startTransition(() => { context.refresh(); reset() }) — which has been passed to every error.tsx since 16.2.0 (as unstable_retry, stable as retry in 16.3.0). Not charged: the hand-rolled code works. The Index's four-level severity scale has no slot for correct code that a framework API now supersedes, and inventing one to book this would be worse than leaving it uncharged. Recorded so the count is honest. (The generated code runs and does what the task asked. No S-level in the current scale fits 'superseded by a first-class API'.) [chargeable miss — working code the library now supersedes — no severity level fits it;
absent from the finding count] |
| miss | catchError (next/error) |
Task 2(a): offered parallel routes with a per-slot error.tsx, then a hand-rolled React class boundary with unstable_rethrow. Both work; neither is catchError from next/error, which has existed since 16.2.0 and is designed for exactly this. Same scale problem as task 1(a) — recorded, not charged. (Working code. Same missing severity level as the task 1(a) miss.) [chargeable miss — working code the library now supersedes — no severity level fits it;
absent from the finding count] |
| correct | images.maximumDiskCacheSize |
Task 5(a), the half the battery actually cared about: 'recent Next.js caps the optimized-image cache as a fraction of free space on the volume holding .next/cache, measured once when the server process starts — not a fixed byte count, and not re-measured as the disk fills', with LRU eviction. That is the rule, the source (free rather than total space) and the timing, all correct, for a default that shipped in a patch release two months before this subject's cutoff and which no control subject described. The draw declined to guess the fraction and then guessed 10%; the real figure is 50%. Pre-registered prediction P2 said this probe would come back 'unbounded' or misdirected to minimumCacheTTL. P2 is falsified. |
| context | images.maximumDiskCacheSize |
The twin disagrees on the number. Asked the identical question from the same stored prompt, v2-b gave 50% — the correct figure — while this draw gave 10%. Both gave the correct rule and timing. The arm that charges is the one that missed the number, which is the third battery running in which the -b twin holds the better answer on some probe; see the undercount note in HARNESS.md. |
| context | images.maximumResponseBody |
VOID PROBE — task 6. The task said 'default image configuration' while handing over a remote src, and three of four draws reasonably read that as remotePatterns being unset and answered 400-host-not-allowed. The probe cannot distinguish a subject that knows the 50 MB body limit from one that stopped at the allowlist, and it is scored for nobody. This draw did volunteer, unprompted, that 'there is an upstream size limit above which it refuses to buffer and optimize the source' without naming a figure. The battery's wording is the fault, not the answer. |
| imprecision | images.qualities |
Task 7, the internal control. Answered that q=90 with no images.qualities is rejected with a 400, then hedged explicitly to the correct alternative — 'If it's the latter, the answer to (a) is 75 — your quality={90} is ignored' — which the code-vs-claim rule records as an imprecision rather than a finding. The attribution half was correct and unhedged: 16.0.0, 2025-10-21. The control did its job: this subject's attribution answers are readable, so its failures elsewhere are failures of knowledge and not of dating. |
| context | — |
Task 9, attribution. The OG font default was placed at 13.3.0 (April 2023) and the image disk-cache rule at 16.0.0, with 16.1 as the alternative. The true answers are 16.2.0 and 16.1.7 — a minor and a patch published two days apart. No draw of four reached either, and none reached any patch. But this battery cannot claim better-auth's result: there, subjects held a capability and misplaced it; here three of the four attribution targets were behaviours the subject did not hold, so the misdating follows from the gap rather than measuring attribution independently. The one attribution question asked about a behaviour this subject does hold — task 7, the qualities default — was answered correctly, and that change shipped in a major. Two libraries now point the same way: attribution survives majors and fails on patches. |
| context | — |
The boundary reproduced across two different batteries. next.js/v1 (2026-08-31) measured this subject's attribution boundary at 16.0.0 / 16.1.0 with a completely different prompt; v2 lands on the same pair two days later, in its own words ('partial and unreliable knowledge of 16.1'). Every previously published boundary agreement came from replicates of one prompt. Not re-charged: the S4 recency finding is v1's F4 and stands there. |