Run next.js--claude-sonnet-5--v2-c--2026-09-02
A below-floor control arm. This subject's next.js attribution stops at 15.0.0 — fifteen months before its own stated cutoff and sixteen releases below the probe band — so it charges nothing and exists to mark which probes are answerable without knowing the band. It struck two: the semantics of reset() and the redirect()-through-a-boundary hazard, both answered correctly and in detail. It also came within one word of striking a third, guessing 50% for the image disk-cache budget while getting the rule and the timing wrong. Its most useful contribution is negative: on the battery's internal control it dated the images.qualities default change to 15.3, not 16.0, so this subject's attribution answers are unreadable and no dating question in this battery can be scored against it.
| Subject | Claude Sonnet 5 claude-sonnet-5, Anthropic |
|---|---|
| Invoked as | Agent tool, model alias "sonnet" |
| Cutoff the model states | 2026-01 |
| Newest next.js release it could place | 15.0.0 · 2024-10-21 (~15 month lag) |
| Oldest next.js release it could not place | 15.1.0 · 2024-12-10 (so this run brackets the subject’s boundary to 2024-10-21 – 2024-12-10) |
| In its own words | "The most recent version name I have some awareness of is in the Next.js 15.x line (I have a vague, unreliable sense that development/canary work toward a Next.js 16 may have been underway, but I can't describe its actual contents) ... the most recent release whose contents I can actually describe is Next.js 15.0, released October 2024" |
| Library at test time | next.js 16.3.4 (npm), verified 2026-09-02 |
| Battery | next.js/v2-c · 9 tasks, 2 direct questions · probe window 16.1.5 to 16.2.0 |
| Tool uses during test | 0 (a run with any tool use is void — we measure training knowledge, not retrieval) |
| Tested | 2026-09-02 |
| Findings | 0, of which 0 chargeable |
None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.
Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.
| Kind | API | Note |
|---|---|---|
| context | error.tsx retry prop |
DERIVABLE — task 1(b). Answered correctly and unprompted from fifteen months below the band: 'reset() is purely a React error-boundary state reset ... does not talk to the network, does not invalidate the Next.js Router Cache, and does not re-invoke the Server Component ... the database query does not run again'. No pass on this probe by any arm is reported as knowledge. This is pre-registered prediction P4, confirmed. |
| context | catchError (next/error) |
DERIVABLE — task 2(b). Described the NEXT_REDIRECT digest being swallowed by a hand-rolled boundary and named unstable_rethrow as the fix, from below the floor. Prediction P5, confirmed. |
| context | images.maximumDiskCacheSize |
PARTIALLY DERIVABLE — task 5(a). Reached the figure without the rule: '50% of the volume's total disk space, checked at the point a new optimized variant is about to be written (using a disk-space check, not computed once at server boot)'. The true rule is 50% of free space computed once at startup, and the task explicitly asked when it is computed, so this is not a pass. But the number is evidently reachable from below the floor, which downgrades the number half of the test-arm's answer. The rule-and-timing half is not derivable: this draw asserted the opposite of it, and the other control denied any such cap exists. |
| miss | ImageResponse default font |
Task 3: Noto Sans, at ~75% confidence. Four draws of four named Noto Sans. Not chargeable — 16.2.0 is fourteen months above this subject's attribution boundary and two months above its stated cutoff. |
| miss | Link transitionTypes |
Task 4: asserted transitionTypes is not a recognised prop, predicted the React unknown-attribute warning in both routers. Same wrong answer as every other draw. Not chargeable — above this subject's cutoff. |
| miss | experimental.clientSegmentCache |
Task 8: experimental.clientSegmentCache: false, at ~40% confidence on the key name. The key was removed in 16.0.3 on 2025-11-13 — two months INSIDE this subject's stated cutoff, and the failure is therefore chargeable in principle. It is not charged, because a control arm charges nothing. Queued: it wants a real battery against this subject, alongside Fable 5, which made the identical recommendation. [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing; absent from the finding count] |
| miss | images.qualities |
Task 7, the internal control, and the reason this arm's dating answers cannot be read. Got the images.qualities default right — '[75] ... it would fail at request time' at ~45% confidence — and then dated the change to 'Next.js 15.3.0, roughly Q2 2025', flagged as a guess. The real change is 16.0.0, a major this subject cannot describe at all. The internal control was written so the attribution question could come back untestable rather than merely supported; for this subject it came back untestable, which is the outcome it was there to make visible. |
| context | images.maximumDiskCacheSize |
Task 5(b): the only draw of four that refused to name a key rather than guess one — 'I'm not going to fabricate a config key that doesn't exist' — and offered filesystem quotas and cron pruning instead. The answer is wrong (images.maximumDiskCacheSize exists, four minors above its boundary) but the behaviour is the one the Index wants from a model that does not know: the two draws that guessed a key produced config that exits the build. |
| context | — | Boundary reproduced exactly. next.js/v1 measured this subject at 15.0.0 / 15.1.0 on 2026-08-31 with an entirely different prompt; v2 lands on the same pair. Not re-charged — the S4 recency finding belongs to v1. Worth recording that this subject's next.js knowledge stops a full year below Fable 5's, on an identical stated cutoff. |
Battery specification: prompts/nextjs.md in the studio repo.
Every finding above also carries its own citation.