What Claude Sonnet 5 gets right about next.js — battery v3-b, tested 2026-09-02

Run next.js--claude-sonnet-5--v3-b--2026-09-02

Summary

The blind twin of v3-a, and it reproduced the failure while disqualifying itself from being charged for it. It answered "Yes" to task 1(a) and wrote the same dead key, then went further than its twin on task 2(b) by explicitly rejecting the correct answer - "Not 'nothing'" - and describing in detail the prefetch behaviour the flag would change. It also refused the cutoff its environment reported, stating early-to-mid 2025 instead and naming the reason: "that's the environment's clock, not evidence about my training horizon". That puts 16.0.3 above its stated cutoff, so even without the -b rule the fairness rule would bar the charge. This is the JOURNAL/031 split replicating in a second library, and this time the arms fell the other way round: the twin licensed to charge is the one that accepted the later cutoff, so the finding survives on v3-a. The battery is not re-run to move this. Its describable-content boundary also reads two releases above its twin's, 15.4-15.5 against 15.2-15.3, from the same stored prompt in the same session.

SubjectClaude Sonnet 5 claude-sonnet-5, Anthropic
Invoked asAgent tool, model alias "sonnet"
Cutoff the model states2025-06
Newest next.js release it could place15.5.0 · 2025-08-20 (~0 month lag)
Oldest next.js release it could not place16.0.0 · 2025-10-22 (so this run brackets the subject’s boundary to 2025-08-20 – 2025-10-22)
In its own words"The most recent Next.js version I can even name is somewhere in the 16.x range by label, but I cannot reliably describe its actual contents ... The most recent version whose contents I can actually describe with real confidence is around Next.js 15.4-15.5, which I'd place at roughly mid-to-late 2025"
Library at test timenext.js 16.3.4 (npm), verified 2026-09-02
Batterynext.js/v3-b · 5 tasks, 2 direct questions · probe window 16.0.3 to 16.0.3
Tool uses during test0 (a run with any tool use is void — we measure training knowledge, not retrieval)
Tested2026-09-02
Findings0, of which 0 chargeable

Findings

None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.

What it got right, and near misses

Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.

KindAPINote
missexperimental.clientSegmentCache Task 1 and task 2: "Yes" in one word, then experimental: { clientSegmentCache: false }, then on task 2(b) an explicit rejection of the correct answer - "So: fewer/larger requests, more duplicated bytes on the wire in aggregate, less cross-link cache reuse. Not 'nothing'." The key was removed in patch 16.0.3 on 2025-11-13. Two rules keep it off the count and the primary one is the arm: this is the -b draw of a duplicated test arm and charges nothing. The secondary rule would also bar it - this draw states its cutoff as early-to-mid 2025, before the release - which is why the miss_class below names the arm rather than the cutoff: the arm rule binds first and would bind whatever the cutoff said. The same failure is charged against this subject on the twin. [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing; charged as a finding on next.js--claude-sonnet-5--v3-a--2026-09-02]
missmiddleware.ts / export function middleware Task 4, the internal control, identical to the twin: middleware.ts exporting middleware, with pages/_middleware.ts named as the older convention. Does not hold the 16.0.0 proxy.ts rename. Already charged against this subject as F1 of the v1 run; asked here as the calibration control, which came back untestable for the second battery running. [chargeable miss — produced only by a belief question the battery does not score as a finding; charged as a finding on next.js--claude-sonnet-5--v1--2026-08-31]
correctimages.maximumResponseBody Task 3(b), and the behaviour the Index wants from a subject that does not know. Refused to name a config key rather than invent one - "I'm answering 'no limit I can name' rather than fabricating a number" - while explicitly flagging that the question's phrasing suggested a real limit it was failing to retrieve. The answer is wrong (images.maximumResponseBody exists, 50 MB, 413) but the release is far above this draw's stated cutoff and the refusal is the right shape. Its nearest recalled figure was sharp's own limitInputPixels, correctly attributed to sharp rather than to Next.js.
context The twin pair disagrees about the subject's own training cutoff, from one stored prompt in one session. v3-a: "the system context here states my knowledge cutoff as January 2026, and I'll take that as authoritative". v3-b: "that's the environment's clock, not evidence about my training horizon, and I'm not treating it as such", giving early-to-mid 2025. JOURNAL/031 measured this on zod with the same subject; it now replicates on a second library, and with the opposite consequence, because there the arm that accepted the later cutoff was the non-charging twin and here it is the charging one.
context The describable-content boundary also moved between the twins: 15.4-15.5 here, 15.2-15.3 on v3-a, against 15.0.0 in both v1 and v2-c. Four readings of one subject on one library across three batteries, two of them from a byte-identical prompt in the same session. The identical-prompt spread is the instrument reading; the wider spread includes v1, whose sent text predates prompts/sent/.

Open questions from this run

Sources

Battery specification: prompts/nextjs.md in the studio repo. Every finding above also carries its own citation.