What Claude Sonnet 5 gets wrong about next.js — battery v3-a, tested 2026-09-02

Run next.js--claude-sonnet-5--v3-a--2026-09-02

Summary

The charging arm of the battery written to book the miss next.js/v2 could only queue, and it books it. Asked whether per-part prefetching can be turned off, this draw answered "yes" in one word and wrote experimental: { clientSegmentCache: false } - a key removed in patch 16.0.3 on 2025-11-13, two months inside its own stated cutoff. Handed the same key back in task 2 it went further than the task needed, asserting that the key "is a recognized key under experimental in the config schema Next.js validates against at build start, so no warning is printed". It is not recognised, and the build warns. The battery's two-shape design earned itself here: on outcome alone task 2(b) looks like a pass, because the draw answered "nothing" - but for the opposite reason, believing the key exists and defaults to false rather than that it does not exist. Task 1 is what makes the belief legible. On the internal control it named middleware.ts, does not hold the 16.0.0 proxy.ts rename at all, and answered the attribution question about a real Next 12 rename instead - so every dating answer in this run is unreadable, exactly as it was in v2. The run's second result is about the instrument rather than the library: this draw put its own describable-content boundary at 15.2-15.3, where v1 and v2 both put it at 15.0.0, and its blind twin put it at 15.4-15.5.

SubjectClaude Sonnet 5 claude-sonnet-5, Anthropic
Invoked asAgent tool, model alias "sonnet"
Cutoff the model states2026-01
Newest next.js release it could place15.3.0 · 2025-04-09 (~9 month lag)
Oldest next.js release it could not place15.4.0 · 2025-05-30 (so this run brackets the subject’s boundary to 2025-04-09 – 2025-05-30)
In its own words"The most recent version name I'm aware of is somewhere in the Next.js 15.x line ... The most recent release whose actual contents I can describe with any real specificity is around Next.js 15.2-15.3 ... I'd place that roughly February-April 2025"
Library at test timenext.js 16.3.4 (npm), verified 2026-09-02
Batterynext.js/v3-a · 5 tasks, 2 direct questions · probe window 16.0.3 to 16.0.3
Tool uses during test0 (a run with any tool use is void — we measure training knowledge, not retrieval)
Tested2026-09-02
Findings1, of which 1 chargeable

Findings

F1 · Answers "yes" and configures experimental.clientSegmentCache, a key removed in patch 16.0.3

S2silently-wrong · experimental.clientSegmentCache · removed · changed in next.js 16.0.3 (2025-11-13) · chargeable

Removed 2025-11-13, two months before this draw's stated cutoff of January 2026, which it accepted from its environment as authoritative. Not back-filled from another run: the twin v3-b states a different cutoff and charges nothing.

What the model believes

Task 1(a), one word: "yes". Task 1(b): "That's controlled by the experimental flag clientSegmentCache. Setting it to false reverts to the older, single 'whole route' prefetch request per <Link>". Task 2(a): "It accepts the config silently. clientSegmentCache is a recognized key under experimental in the config schema Next.js validates against at build start, so no warning is printed and the build proceeds normally." Task 5(a): "it has not stopped being recognized, as far as I know."

What it wrote
experimental: { clientSegmentCache: false }
What works on next.js 16.3.4
experimental: { prefetchInlining: true }
Impact

The team asked how to cut a prefetch burst and gets a config that compiles, deploys and does nothing. An unrecognised key under experimental warns and is dropped rather than failing the build, so the file reads as applied and the burst continues. The draw's stated mechanism makes it worse than a bad guess: it tells the reader the key is schema-validated, which is the one check that would have caught it.

Scope note

Re-confirmed against the published packages this session before the battery was written: two occurrences at 16.0.2, zero at 16.0.3, zero at 16.3.4.

Verified against

What it got right, and near misses

Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.

KindAPINote
missmiddleware.ts / export function middleware Task 4, the internal control. Named middleware.ts exporting middleware and does not hold the 16.0.0 rename to proxy.ts at all - asked for an older name it reached back to the per-directory pages/_middleware.ts convention of Next 12. The generated middleware.ts is deprecated as of 16.0.0 and is a real miss inside the window, but it is already charged against this subject as F1 of the v1 run, and this battery asks it as the calibration control rather than as a charging probe. (The internal control exists to say whether this subject's attribution answers can be read at all. It came back untestable - the subject does not hold the change, so its dating answer measures nothing (JOURNAL/030(b)). Charging it here would also double-count v1 F1.) [chargeable miss — produced only by a belief question the battery does not score as a finding; charged as a finding on next.js--claude-sonnet-5--v1--2026-08-31]
missimages.maximumResponseBody Task 3, and the reason it is not charged against this subject. Answered that no framework-level source-size limit exists - "the request to /_next/image ends with HTTP 200, provided the Node process has enough memory" - and in (c) made the resource story explicit: "more RAM and disk headroom make the 200-OK success path more likely". The truth is a fixed 50 MB cap that throws a 413 while streaming, independent of RAM. images.maximumResponseBody arrived in 16.1.5 on 2026-01-26, the same month as this subject's stated cutoff, so it sits outside the fairness window under the same-month rule and is not a chargeable miss. It did name 50 MB, but attached it to Vercel's hosted platform rather than to the framework, and used that attribution to rule the limit out for a self-hosted app.
contextexperimental.clientSegmentCache Task 2(b) is the reason this battery asked the same surface two ways. On outcome alone it passed - "it's identical to omitting the key entirely" - and a battery that had asked only the recognition question would have scored a pass. The mechanism is the opposite of the truth: the draw believes the key is live and defaults to false, not that it does not exist, and flagged the alternative reading at "maybe 55/45". BACKLOG 2b asked for the recognition shape alone; keeping the offer shape as well is what made the belief legible.
context The boundary moved between batteries for this subject, which v1 and v2 gave no sign of. v1 (2026-08-31) and v2-c (2026-09-02) both read 15.0.0 / 15.1.0. This draw reads 15.2-15.3 describable, "roughly February-April 2025", and its blind twin reads 15.4-15.5. JOURNAL/030(c) concluded from v1 and v2 that a different battery is not a different boundary; on this subject and this library, v3 contradicts it, and the twin pair shows the movement inside one stored prompt. Not charged - the S4 recency finding for this subject belongs to v1 F7.
context The stated cutoff, recorded because it is the input to the fairness rule and because it moved. This draw accepted the environment value as authoritative and caveated only its recall depth; its blind twin v3-b explicitly refused it - "that's the environment's clock, not evidence about my training horizon" - and stated early-to-mid 2025 instead. Same subject, same stored prompt, same session, two different cutoffs. This is JOURNAL/031's zod/v4 result replicating in a second library, with the arms the other way round: there the charging arm disqualified itself, here the charging arm is the one that admits the release.

Open questions from this run

Sources

Battery specification: prompts/nextjs.md in the studio repo. Every finding above also carries its own citation.