# The Stale Priors Index — full text export # Generated from data/index.json. Canonical JSON: https://stalepriors.com/data/index.json # Last updated 2026-09-06. Runs: 152. Findings: 162 (155 chargeable). ## What this is A dataset of reproduced failures of AI coding models on fast-moving libraries. Each finding records what the model believed, the code it wrote, the code that works, and the primary source proving the difference. A finding is 'chargeable' only if the library change predates the model's own stated training cutoff. Runs are never deleted; a retest supersedes. ## Headline measurement: version attribution lags the stated cutoff - Claude Fable 5 on better-auth: states cutoff 2026-01; version attribution stops at 1.3.0 (2025-07-19); lag ~6 months. - Claude Fable 5 on better-auth: states cutoff 2026-01; version attribution stops at 1.3.0 (2025-07-19); lag ~6 months. - Claude Fable 5 on better-auth: states cutoff 2026-01; version attribution stops at 1.3.0 (2025-07-19); lag ~6 months. - Claude Fable 5 on better-auth: states cutoff 2026-01; version attribution stops at 1.3.0 (2025-07-19); lag ~18 months. - Claude Fable 5 on better-auth: states cutoff 2026-01; version attribution stops at 1.3.0 (2025-07-19); lag ~18 months. - Claude Fable 5 on better-auth: states cutoff 2026-01; version attribution stops at 1.3.0 (2025-07-19); lag ~18 months. - Claude Fable 5 on better-auth: states cutoff 2026-01; version attribution stops at 1.3.0 (2025-07-19); lag ~6 months. - Claude Fable 5 on langchain: states cutoff 2026-01; version attribution stops at 1.0.0 (2025-10-17); lag ~3 months. - Claude Fable 5 on next.js: states cutoff 2026-01; version attribution stops at 16.0.0 (2025-10-22); lag ~3 months. - Claude Fable 5 on next.js: states cutoff 2026-01; version attribution stops at 16.0.0 (2025-10-22); lag ~3 months. - Claude Fable 5 on next.js: states cutoff 2026-01; version attribution stops at 16.0.0 (2025-10-22); lag ~3 months. - Claude Fable 5 on next.js: states cutoff 2026-01; version attribution stops at 16.0.0 (2025-10-22); lag ~3 months. - Claude Fable 5 on prisma: states cutoff 2026-01; version attribution stops at 6.7.0 (2025-04-29); lag ~8 months. - Claude Fable 5 on prisma: states cutoff 2026-01; version attribution stops at 6.7.0 (2025-04-29); lag ~8 months. - Claude Fable 5 on tailwindcss: states cutoff 2026-01; version attribution stops at 4.1.0 (2025-04-01); lag ~9 months. - Claude Fable 5 on tailwindcss: states cutoff 2026-01; version attribution stops at 4.1.0 (2025-04-01); lag ~9 months. - Claude Fable 5 on valibot: states cutoff 2026-01; version attribution stops at 1.1.0 (2025-05-06); lag ~8 months. - Claude Fable 5 on valibot: states cutoff 2026-01; version attribution stops at 1.1.0 (2025-05-06); lag ~8 months. - Claude Fable 5 on valibot: states cutoff 2026-01; version attribution stops at 1.1.0 (2025-05-06); lag ~8 months. - Claude Fable 5 on zod: states cutoff 2026-01; version attribution stops at 4.1.0 (2025-08-23); lag ~5 months. - Claude Fable 5 on zod: states cutoff 2026-01; version attribution stops at 4.1.0 (2025-08-23); lag ~5 months. - Claude Fable 5 on zod: states cutoff 2026-01; version attribution stops at 4.1.0 (2025-08-23); lag ~5 months. - Claude Fable 5 on zod: states cutoff 2026-01; version attribution stops at 4.1.0 (2025-08-23); lag ~5 months. - Claude Fable 5.1 on better-auth: states cutoff 2026-06; version attribution stops at 1.3.0 (2025-07-19); lag ~11 months. - Claude Fable 5.1 on better-auth: states cutoff 2026-06; version attribution stops at 1.4.0 (2025-11-22); lag ~7 months. - Claude Fable 5.1 on better-auth: states cutoff 2026-06; version attribution stops at 1.3.0 (2025-07-19); lag ~11 months. - Claude Fable 5.1 on langchain: states cutoff 2026-06; version attribution stops at 1.0.0 (2025-10-17); lag ~8 months. - Claude Fable 5.1 on langchain: states cutoff 2026-06; version attribution stops at 1.1.0 (2025-11-24); lag ~6 months. - Claude Fable 5.1 on langchain: states cutoff 2026-06; version attribution stops at 1.1.0 (2025-11-24); lag ~6 months. - Claude Fable 5.1 on langchain: states cutoff 2026-06; version attribution stops at 1.2.0 (2025-12-15); lag ~6 months. - Claude Fable 5.1 on next.js: states cutoff 2026-06; version attribution stops at 16.1.0 (2025-12-18); lag ~6 months. - Claude Fable 5.1 on next.js: states cutoff 2026-06; version attribution stops at 16.1.0 (2025-12-18); lag ~6 months. - Claude Fable 5.1 on next.js: states cutoff 2026-06; version attribution stops at 16.1.0 (2025-12-18); lag ~6 months. - Claude Fable 5.1 on next.js: states cutoff 2026-06; version attribution stops at 16.1.0 (2025-12-18); lag ~6 months. - Claude Fable 5.1 on prisma: states cutoff 2026-06; version attribution stops at 7.0.0 (2025-11-19); lag ~7 months. - Claude Fable 5.1 on prisma: states cutoff 2026-06; version attribution stops at 7.0.0 (2025-11-19); lag ~7 months. - Claude Fable 5.1 on prisma: states cutoff 2026-06; version attribution stops at 7.0.0 (2025-11-19); lag ~7 months. - Claude Fable 5.1 on prisma: states cutoff 2026-06; version attribution stops at 7.2.0 (2025-12-17); lag ~6 months. - Claude Fable 5.1 on tailwindcss: states cutoff 2026-06; version attribution stops at 4.1.0 (2025-04-01); lag ~14 months. - Claude Fable 5.1 on tailwindcss: states cutoff 2026-06; version attribution stops at 4.1.0 (2025-04-01); lag ~14 months. - Claude Fable 5.1 on valibot: states cutoff 2026-06; version attribution stops at 1.1.0 (2025-05-06); lag ~13 months. - Claude Fable 5.1 on valibot: states cutoff 2026-06; version attribution stops at 1.1.0 (2025-05-06); lag ~13 months. - Claude Fable 5.1 on zod: states cutoff 2026-06; version attribution stops at 4.1.0 (2025-08-23); lag ~10 months. - Claude Haiku 4.5 on langchain: states cutoff 2025-02; version attribution stops at 0.3.0 (2024-09-13); lag ~5 months. - Claude Haiku 4.5 on langchain: states cutoff 2025-02; version attribution stops at 0.2.0 (2024-05-20); lag ~9 months. - Claude Opus 5 on better-auth: states cutoff 2026-05; version attribution stops at 1.3.0 (2025-07-19); lag ~10 months. - Claude Opus 5 on better-auth: states cutoff 2026-05; version attribution stops at 1.3.0 (2025-07-19); lag ~10 months. - Claude Opus 5 on better-auth: states cutoff 2026-05; version attribution stops at 1.3.0 (2025-07-19); lag ~10 months. - Claude Opus 5 on better-auth: states cutoff 2026-05; version attribution stops at 1.3.0 (2025-07-19); lag ~10 months. - Claude Opus 5 on better-auth: states cutoff 2026-05; version attribution stops at 1.3.0 (2025-07-19); lag ~10 months. - Claude Opus 5 on better-auth: states cutoff 2026-05; version attribution stops at 1.2.0 (2025-03-01); lag ~22 months. - Claude Opus 5 on better-auth: states cutoff 2026-05; version attribution stops at 1.2.0 (2025-03-01); lag ~22 months. - Claude Opus 5 on better-auth: states cutoff 2026-05; version attribution stops at 1.2.0 (2025-03-01); lag ~22 months. - Claude Opus 5 on better-auth: states cutoff 2026-05; version attribution stops at 1.2.0 (2025-03-01); lag ~22 months. - Claude Opus 5 on better-auth: states cutoff 2026-05; version attribution stops at 1.2.0 (2025-03-01); lag ~22 months. - Claude Opus 5 on better-auth: states cutoff 2026-05; version attribution stops at 1.2.0 (2025-03-01); lag ~22 months. - Claude Opus 5 on better-auth: states cutoff 2026-05; version attribution stops at 1.2.0 (2025-03-01); lag ~26 months. - Claude Opus 5 on better-auth: states cutoff 2026-05; version attribution stops at 1.2.0 (2025-03-01); lag ~26 months. - Claude Opus 5 on langchain: states cutoff 2026-05; version attribution stops at 1.0.0 (2025-10-17); lag ~7 months. - Claude Opus 5 on langchain: states cutoff 2026-05; version attribution stops at 1.0.0 (2025-10-17); lag ~7 months. - Claude Opus 5 on langchain: states cutoff 2026-05; version attribution stops at 1.0.0 (2025-10-17); lag ~7 months. - Claude Opus 5 on langchain: states cutoff 2026-05; version attribution stops at 1.0.0 (2025-10-17); lag ~7 months. - Claude Opus 5 on next.js: states cutoff 2026-05; version attribution stops at 16.0.0 (2025-10-22); lag ~7 months. - Claude Opus 5 on next.js: states cutoff 2026-05; version attribution stops at 16.0.0 (2025-10-22); lag ~7 months. - Claude Opus 5 on next.js: states cutoff 2026-05; version attribution stops at 16.0.0 (2025-10-22); lag ~7 months. - Claude Opus 5 on next.js: states cutoff 2026-05; version attribution stops at 16.0.0 (2025-10-22); lag ~7 months. - Claude Opus 5 on next.js: states cutoff 2026-05; version attribution stops at 16.0.0 (2025-10-22); lag ~7 months. - Claude Opus 5 on next.js: states cutoff 2026-05; version attribution stops at 16.0.0 (2025-10-22); lag ~7 months. - Claude Opus 5 on next.js: states cutoff 2026-05; version attribution stops at 16.0.0 (2025-10-22); lag ~7 months. - Claude Opus 5 on prisma: states cutoff 2026-05; version attribution stops at 7.0.0 (2025-11-19); lag ~5 months. - Claude Opus 5 on prisma: states cutoff 2026-05; version attribution stops at 7.0.0 (2025-11-19); lag ~5 months. - Claude Opus 5 on prisma: states cutoff 2026-05; version attribution stops at 6.7.0 (2025-04-29); lag ~12 months. - Claude Opus 5 on prisma: states cutoff 2026-05; version attribution stops at 6.7.0 (2025-04-29); lag ~12 months. - Claude Opus 5 on prisma: states cutoff 2026-05; version attribution stops at 7.0.0 (2025-11-19); lag ~5 months. - Claude Opus 5 on prisma: states cutoff 2026-05; version attribution stops at 6.8.0 (2025-05-15); lag ~12 months. - Claude Opus 5 on prisma: states cutoff 2026-05; version attribution stops at 6.10.0 (2025-06-17); lag ~11 months. - Claude Opus 5 on prisma: states cutoff 2026-05; version attribution stops at 7.0.0 (2025-11-19); lag ~6 months. - Claude Opus 5 on prisma: states cutoff 2026-05; version attribution stops at 6.15.0 (2025-08-27); lag ~9 months. - Claude Opus 5 on prisma: states cutoff 2026-05; version attribution stops at 7.0.0 (2025-11-19); lag ~6 months. - Claude Opus 5 on prisma: states cutoff 2026-05; version attribution stops at 6.7.0 (2025-04-29); lag ~13 months. - Claude Opus 5 on prisma: states cutoff 2026-05; version attribution stops at 7.0.0 (2025-11-19); lag ~6 months. - Claude Opus 5 on tailwindcss: states cutoff 2026-05; version attribution stops at 4.1.0 (2025-04-01); lag ~13 months. - Claude Opus 5 on tailwindcss: states cutoff 2026-05; version attribution stops at 4.1.0 (2025-04-01); lag ~13 months. - Claude Opus 5 on tailwindcss: states cutoff 2026-05; version attribution stops at 4.1.0 (2025-04-01); lag ~13 months. - Claude Opus 5 on tailwindcss: states cutoff 2026-05; version attribution stops at 4.1.0 (2025-04-01); lag ~13 months. - Claude Opus 5 on valibot: states cutoff 2026-05; version attribution stops at 1.1.0 (2025-05-06); lag ~12 months. - Claude Opus 5 on valibot: states cutoff 2026-05; version attribution stops at 1.0.0 (2025-03-19); lag ~14 months. - Claude Opus 5 on valibot: states cutoff 2026-05; version attribution stops at 1.0.0 (2025-03-19); lag ~14 months. - Claude Opus 5 on valibot: states cutoff 2026-05; version attribution stops at 1.0.0 (2025-03-19); lag ~14 months. - Claude Opus 5 on valibot: states cutoff 2026-05; version attribution stops at 1.1.0 (2025-05-06); lag ~12 months. - Claude Opus 5 on valibot: states cutoff 2026-05; version attribution stops at 1.1.0 (2025-05-06); lag ~12 months. - Claude Opus 5 on valibot: states cutoff 2026-05; version attribution stops at 1.1.0 (2025-05-06); lag ~12 months. - Claude Opus 5 on valibot: states cutoff 2026-05; version attribution stops at 1.1.0 (2025-05-06); lag ~12 months. - Claude Opus 5 on zod: states cutoff 2026-05; version attribution stops at 4.1.0 (2025-08-23); lag ~9 months. - Claude Opus 5 on zod: states cutoff 2026-05; version attribution stops at 4.1.0 (2025-08-23); lag ~9 months. - Claude Opus 5 on zod: states cutoff 2026-05; version attribution stops at 4.1.0 (2025-08-23); lag ~9 months. - Claude Opus 5 on zod: states cutoff 2026-05; version attribution stops at 4.1.0 (2025-08-23); lag ~9 months. - Claude Opus 5 on zod: states cutoff 2026-05; version attribution stops at 4.1.0 (2025-08-23); lag ~9 months. - Claude Sonnet 5 on better-auth: states cutoff 2026-01; version attribution stops at 1.0.0 (2024-11-23); lag ~13 months. - Claude Sonnet 5 on better-auth: states cutoff 2026-01; version attribution stops at 1.0.0 (2024-11-23); lag ~13 months. - Claude Sonnet 5 on better-auth: states cutoff 2026-01; version attribution stops at 1.2.0 (2025-03-01); lag ~18 months. - Claude Sonnet 5 on better-auth: states cutoff 2026-01; version attribution stops at 1.0.0 (2024-11-23); lag ~13 months. - Claude Sonnet 5 on better-auth: states cutoff 2026-01; version attribution stops at 1.1.0 (2024-12-01); lag ~13 months. - Claude Sonnet 5 on better-auth: states cutoff 2026-01; version attribution stops at 1.0.0 (2024-11-23); lag ~14 months. - Claude Sonnet 5 on langchain: states cutoff 2026-01; version attribution stops at 0.3.0 (2024-09-13); lag ~16 months. - Claude Sonnet 5 on langchain: states cutoff 2026-01; version attribution stops at 1.0.0 (2025-10-17); lag ~3 months. - Claude Sonnet 5 on langchain: states cutoff 2026-01; version attribution stops at 0.3.0 (2024-09-13); lag ~16 months. - Claude Sonnet 5 on langchain: states cutoff 2026-01; version attribution stops at 1.0.0 (2025-10-17); lag ~3 months. - Claude Sonnet 5 on langchain: states cutoff 2026-01; version attribution stops at 0.3.0 (2024-09-13); lag ~16 months. - Claude Sonnet 5 on langchain: states cutoff 2026-01; version attribution stops at 0.3.0 (2024-09-13); lag ~16 months. - Claude Sonnet 5 on next.js: states cutoff 2026-01; version attribution stops at 15.0.0 (2024-10-21); lag ~15 months. - Claude Sonnet 5 on next.js: states cutoff 2026-01; version attribution stops at 15.0.0 (2024-10-21); lag ~15 months. - Claude Sonnet 5 on next.js: states cutoff 2026-01; version attribution stops at 15.3.0 (2025-04-09); lag ~9 months. - Claude Sonnet 5 on next.js: states cutoff 2025-06; version attribution stops at 15.5.0 (2025-08-20); lag ~0 months. - Claude Sonnet 5 on prisma: states cutoff 2026-01; version attribution stops at 6.0.0 (2024-11-28); lag ~14 months. - Claude Sonnet 5 on prisma: states cutoff 2026-01; version attribution stops at 6.0.0 (2024-11-28); lag ~14 months. - Claude Sonnet 5 on prisma: states cutoff 2026-01; version attribution stops at 6.0.0 (2024-11-28); lag ~14 months. - Claude Sonnet 5 on prisma: states cutoff 2026-01; version attribution stops at 6.0.0 (2024-11-28); lag ~14 months. - Claude Sonnet 5 on tailwindcss: states cutoff 2026-01; version attribution stops at 4.1.0 (2025-04-01); lag ~9 months. - Claude Sonnet 5 on tailwindcss: states cutoff 2026-01; version attribution stops at 4.1.0 (2025-04-01); lag ~9 months. - Claude Sonnet 5 on tailwindcss: states cutoff 2026-01; version attribution stops at 4.1.0 (2025-04-01); lag ~9 months. - Claude Sonnet 5 on valibot: states cutoff 2026-01; version attribution stops at 1.0.0 (2025-03-19); lag ~9.5 months. - Claude Sonnet 5 on valibot: states cutoff 2026-01; version attribution stops at 1.0.0 (2025-03-19); lag ~10 months. - Claude Sonnet 5 on valibot: states cutoff 2026-01; version attribution stops at 1.0.0 (2025-03-19); lag ~10 months. - Claude Sonnet 5 on valibot: states cutoff 2026-01; version attribution stops at 1.0.0 (2025-03-19); lag ~10 months. - Claude Sonnet 5 on zod: states cutoff 2026-01; version attribution stops at 4.0.0 (2025-07-10); lag ~6 months. - Claude Sonnet 5 on zod: states cutoff 2026-01; version attribution stops at 4.0.0 (2025-07-10); lag ~6 months. ## Attribution boundary: is it a single date? Each run brackets a subject's attribution boundary on one library: at or after the newest release whose contents it could correctly attribute to that release, strictly before the oldest release it could not. If a model has one attribution boundary, all its brackets contain it and they intersect. Computed, not asserted. This measures which release carried a feature, NOT whether the model knows the feature. The two come apart: on better-auth one subject placed its boundary at 1.0.0 while correctly using APIs introduced in 1.1.0 and 1.2.0. Mis-attributions observed so far always run EARLY -- to a release older than the one that shipped the feature. - Claude Fable 5 (states cutoff 2026-01): brackets from 7 libraries do NOT intersect — no single date explains this model. It describes next.js 16.0.0 (2025-10-22) but not prisma 6.8.0 (2025-05-15), published 160 days earlier. better-auth: places 1.3.0 (2025-07-19), cannot place 1.4.0 (2025-11-22) — https://stalepriors.com/runs/better-auth--claude-fable-5--v2-a--2026-09-02 better-auth: places 1.3.0 (2025-07-19), cannot place 1.4.0 (2025-11-22) — https://stalepriors.com/runs/better-auth--claude-fable-5--v2-b--2026-09-02 better-auth: places 1.3.0 (2025-07-19), cannot place 1.4.0 (2025-11-22) — https://stalepriors.com/runs/better-auth--claude-fable-5--v3-d--2026-09-03 better-auth: places 1.3.0 (2025-07-19), cannot place 1.4.0 (2025-11-22) — https://stalepriors.com/runs/better-auth--claude-fable-5--v4-e--2026-09-03 better-auth: places 1.3.0 (2025-07-19), cannot place 1.4.0 (2025-11-22) — https://stalepriors.com/runs/better-auth--claude-fable-5--v4-f--2026-09-03 better-auth: places 1.3.0 (2025-07-19), cannot place 1.4.0 (2025-11-22) — https://stalepriors.com/runs/better-auth--claude-fable-5--v1--2026-09-01 langchain: places 1.0.0 (2025-10-17), cannot place 1.1.0 (2025-11-24) — https://stalepriors.com/runs/langchain--claude-fable-5--v1--2026-08-31 next.js: places 16.0.0 (2025-10-22), cannot place 16.1.0 (2025-12-18) — https://stalepriors.com/runs/next.js--claude-fable-5--v2-d--2026-09-02 next.js: places 16.0.0 (2025-10-22), cannot place 16.1.0 (2025-12-18) — https://stalepriors.com/runs/next.js--claude-fable-5--v3-c--2026-09-02 next.js: places 16.0.0 (2025-10-22), cannot place 16.1.0 (2025-12-18) — https://stalepriors.com/runs/next.js--claude-fable-5--v3-d--2026-09-02 next.js: places 16.0.0 (2025-10-22), cannot place 16.1.0 (2025-12-18) — https://stalepriors.com/runs/next.js--claude-fable-5--v1--2026-08-31 prisma: places 6.7.0 (2025-04-29), cannot place 6.8.0 (2025-05-15) — https://stalepriors.com/runs/prisma--claude-fable-5--v2-d--2026-09-05 prisma: places 6.7.0 (2025-04-29), cannot place 6.8.0 (2025-05-15) — https://stalepriors.com/runs/prisma--claude-fable-5--v1--2026-08-31 tailwindcss: places 4.1.0 (2025-04-01), cannot place 4.2.0 (2026-02-18) — https://stalepriors.com/runs/tailwindcss--claude-fable-5--v2-d--2026-09-03 tailwindcss: places 4.1.0 (2025-04-01), cannot place 4.2.0 (2026-02-18) — https://stalepriors.com/runs/tailwindcss--claude-fable-5--v1--2026-08-31 valibot: places 1.1.0 (2025-05-06), cannot place 1.2.0 (2025-11-24) — https://stalepriors.com/runs/valibot--claude-fable-5--v2-e--2026-09-02 valibot: places 1.1.0 (2025-05-06), cannot place 1.2.0 (2025-11-24) — https://stalepriors.com/runs/valibot--claude-fable-5--v2-f--2026-09-02 valibot: places 1.1.0 (2025-05-06), cannot place 1.2.0 (2025-11-24) — https://stalepriors.com/runs/valibot--claude-fable-5--v1--2026-09-01 zod: places 4.1.0 (2025-08-23), cannot place 4.2.0 (2025-12-15) — https://stalepriors.com/runs/zod--claude-fable-5--v2--2026-08-29 zod: places 4.1.0 (2025-08-23), cannot place 4.2.0 (2025-12-15) — https://stalepriors.com/runs/zod--claude-fable-5--v3-d--2026-09-02 zod: places 4.1.0 (2025-08-23), cannot place 4.2.0 (2025-12-15) — https://stalepriors.com/runs/zod--claude-fable-5--v4-a--2026-09-02 zod: places 4.1.0 (2025-08-23), cannot place 4.2.0 (2025-12-15) — https://stalepriors.com/runs/zod--claude-fable-5--v4-b--2026-09-02 - Claude Fable 5.1 (states cutoff 2026-06): brackets from 7 libraries do NOT intersect — no single date explains this model. It describes next.js 16.1.0 (2025-12-18) but not better-auth 1.4.0 (2025-11-22), published 26 days earlier. better-auth (3 draws), langchain (4 draws), prisma (4 draws) also produced brackets that exclude each other within one library; the verdict does not rest on that — no choice of a single draw per library admits any date either. better-auth: places 1.3.0 (2025-07-19), cannot place 1.4.0 (2025-11-22) — https://stalepriors.com/runs/better-auth--claude-fable-5-1--v6--2026-09-06 better-auth: places 1.4.0 (2025-11-22), cannot place 1.5.0 (2026-03-01) — https://stalepriors.com/runs/better-auth--claude-fable-5-1--v7-c--2026-09-06 better-auth: places 1.3.0 (2025-07-19), cannot place 1.4.0 (2025-11-22) — https://stalepriors.com/runs/better-auth--claude-fable-5-1--v7-d--2026-09-06 langchain: places 1.0.0 (2025-10-17), cannot place 1.1.0 (2025-11-24) — https://stalepriors.com/runs/langchain--claude-fable-5-1--v4--2026-09-06 langchain: places 1.1.0 (2025-11-24), cannot place 1.2.0 (2025-12-15) — https://stalepriors.com/runs/langchain--claude-fable-5-1--v5-a--2026-09-06 langchain: places 1.1.0 (2025-11-24), cannot place 1.2.0 (2025-12-15) — https://stalepriors.com/runs/langchain--claude-fable-5-1--v5-b--2026-09-06 langchain: places 1.2.0 (2025-12-15), cannot place 1.3.0 (2026-05-12) — https://stalepriors.com/runs/langchain--claude-fable-5-1--v6-e--2026-09-06 next.js: places 16.1.0 (2025-12-18), cannot place 16.2.0 (2026-03-18) — https://stalepriors.com/runs/next.js--claude-fable-5-1--v4-a--2026-09-05 next.js: places 16.1.0 (2025-12-18), cannot place 16.2.0 (2026-03-18) — https://stalepriors.com/runs/next.js--claude-fable-5-1--v4-b--2026-09-05 next.js: places 16.1.0 (2025-12-18), cannot place 16.2.0 (2026-03-18) — https://stalepriors.com/runs/next.js--claude-fable-5-1--v4-c--2026-09-05 next.js: places 16.1.0 (2025-12-18), cannot place 16.2.0 (2026-03-18) — https://stalepriors.com/runs/next.js--claude-fable-5-1--v4-d--2026-09-05 prisma: places 7.0.0 (2025-11-19), cannot place 7.1.0 (2025-12-03) — https://stalepriors.com/runs/prisma--claude-fable-5-1--v4-c--2026-09-06 prisma: places 7.0.0 (2025-11-19), cannot place 7.1.0 (2025-12-03) — https://stalepriors.com/runs/prisma--claude-fable-5-1--v4-d--2026-09-06 prisma: places 7.0.0 (2025-11-19), cannot place 7.1.0 (2025-12-03) — https://stalepriors.com/runs/prisma--claude-fable-5-1--v5-e--2026-09-06 prisma: places 7.2.0 (2025-12-17), cannot place 7.3.0 (2026-01-21) — https://stalepriors.com/runs/prisma--claude-fable-5-1--v5-f--2026-09-06 tailwindcss: places 4.1.0 (2025-04-01), cannot place 4.2.0 (2026-02-18) — https://stalepriors.com/runs/tailwindcss--claude-fable-5-1--v3-a--2026-09-06 tailwindcss: places 4.1.0 (2025-04-01), cannot place 4.2.0 (2026-02-18) — https://stalepriors.com/runs/tailwindcss--claude-fable-5-1--v3-b--2026-09-06 valibot: places 1.1.0 (2025-05-06), cannot place 1.2.0 (2025-11-24) — https://stalepriors.com/runs/valibot--claude-fable-5-1--v3-a--2026-09-05 valibot: places 1.1.0 (2025-05-06), cannot place 1.2.0 (2025-11-24) — https://stalepriors.com/runs/valibot--claude-fable-5-1--v3-b--2026-09-05 valibot: places 1.1.0 (2025-05-06), cannot place 1.2.0 (2025-11-24) — https://stalepriors.com/runs/valibot--claude-fable-5-1--v4-a--2026-09-06 valibot: places 1.1.0 (2025-05-06), cannot place 1.2.0 (2025-11-24) — https://stalepriors.com/runs/valibot--claude-fable-5-1--v4-b--2026-09-06 zod: places 4.1.0 (2025-08-23), cannot place 4.2.0 (2025-12-15) — https://stalepriors.com/runs/zod--claude-fable-5-1--v5--2026-09-06 - Claude Haiku 4.5: measured on 10 libraries, 1 with a complete bracket — too few to intersect. 8 run(s) establish only one end of the bracket (better-auth, better-auth, better-auth, prisma, prisma, valibot, zod, zod), so they cannot bound a date. - Claude Opus 5 (states cutoff 2026-05): brackets from 7 libraries do NOT intersect — no single date explains this model. It describes prisma 7.0.0 (2025-11-19) but not valibot 1.1.0 (2025-05-06), published 197 days earlier. CAVEAT: that pair is the extreme draw on each side. better-auth (13 draws), prisma (12 draws), valibot (8 draws) produced brackets that exclude each other WITHIN one library, and one choice of a single draw per library still admits 2025-11-19 to 2025-11-22 (3 days) — so this verdict rests on replicate disagreement, not on the libraries disagreeing. better-auth: places 1.3.0 (2025-07-19), cannot place 1.4.0 (2025-11-22) — https://stalepriors.com/runs/better-auth--claude-opus-5--v1r-a--2026-09-01 better-auth: places 1.3.0 (2025-07-19), cannot place 1.4.0 (2025-11-22) — https://stalepriors.com/runs/better-auth--claude-opus-5--v1r-b--2026-09-01 better-auth: places 1.3.0 (2025-07-19), cannot place 1.4.0 (2025-11-22) — https://stalepriors.com/runs/better-auth--claude-opus-5--v2-a--2026-09-02 better-auth: places 1.3.0 (2025-07-19), cannot place 1.4.0 (2025-11-22) — https://stalepriors.com/runs/better-auth--claude-opus-5--v2-b--2026-09-02 better-auth: places 1.2.0 (2025-03-01), cannot place 1.3.0 (2025-07-19) — https://stalepriors.com/runs/better-auth--claude-opus-5--v3-a--2026-09-03 better-auth: places 1.2.0 (2025-03-01), cannot place 1.3.0 (2025-07-19) — https://stalepriors.com/runs/better-auth--claude-opus-5--v3-b--2026-09-03 better-auth: places 1.2.0 (2025-03-01), cannot place 1.3.0 (2025-07-19) — https://stalepriors.com/runs/better-auth--claude-opus-5--v4-c--2026-09-03 better-auth: places 1.2.0 (2025-03-01), cannot place 1.3.0 (2025-07-19) — https://stalepriors.com/runs/better-auth--claude-opus-5--v4-d--2026-09-03 better-auth: places 1.2.0 (2025-03-01), cannot place 1.3.0 (2025-07-19) — https://stalepriors.com/runs/better-auth--claude-opus-5--v5-a--2026-09-03 better-auth: places 1.2.0 (2025-03-01), cannot place 1.3.0 (2025-07-19) — https://stalepriors.com/runs/better-auth--claude-opus-5--v5-b--2026-09-03 better-auth: places 1.2.0 (2025-03-01), cannot place 1.3.0 (2025-07-19) — https://stalepriors.com/runs/better-auth--claude-opus-5--v7-a--2026-09-06 better-auth: places 1.2.0 (2025-03-01), cannot place 1.3.0 (2025-07-19) — https://stalepriors.com/runs/better-auth--claude-opus-5--v7-b--2026-09-06 better-auth: places 1.3.0 (2025-07-19), cannot place 1.4.0 (2025-11-22) — https://stalepriors.com/runs/better-auth--claude-opus-5--v1--2026-09-01 langchain: places 1.0.0 (2025-10-17), cannot place 1.1.0 (2025-11-24) — https://stalepriors.com/runs/langchain--claude-opus-5--v5-c--2026-09-06 langchain: places 1.0.0 (2025-10-17), cannot place 1.1.0 (2025-11-24) — https://stalepriors.com/runs/langchain--claude-opus-5--v6-a--2026-09-06 langchain: places 1.0.0 (2025-10-17), cannot place 1.1.0 (2025-11-24) — https://stalepriors.com/runs/langchain--claude-opus-5--v6-b--2026-09-06 langchain: places 1.0.0 (2025-10-17), cannot place 1.1.0 (2025-11-24) — https://stalepriors.com/runs/langchain--claude-opus-5--v1--2026-08-31 next.js: places 16.0.0 (2025-10-22), cannot place 16.1.0 (2025-12-18) — https://stalepriors.com/runs/next.js--claude-opus-5--v2-a--2026-09-02 next.js: places 16.0.0 (2025-10-22), cannot place 16.1.0 (2025-12-18) — https://stalepriors.com/runs/next.js--claude-opus-5--v2-b--2026-09-02 next.js: places 16.0.0 (2025-10-22), cannot place 16.1.0 (2025-12-18) — https://stalepriors.com/runs/next.js--claude-opus-5--v3-e--2026-09-02 next.js: places 16.0.0 (2025-10-22), cannot place 16.1.0 (2025-12-18) — https://stalepriors.com/runs/next.js--claude-opus-5--v3-f--2026-09-02 next.js: places 16.0.0 (2025-10-22), cannot place 16.1.0 (2025-12-18) — https://stalepriors.com/runs/next.js--claude-opus-5--v4-e--2026-09-05 next.js: places 16.0.0 (2025-10-22), cannot place 16.1.0 (2025-12-18) — https://stalepriors.com/runs/next.js--claude-opus-5--v4-f--2026-09-05 next.js: places 16.0.0 (2025-10-22), cannot place 16.1.0 (2025-12-18) — https://stalepriors.com/runs/next.js--claude-opus-5--v1--2026-08-31 prisma: places 7.0.0 (2025-11-19), cannot place 7.1.0 (2025-12-03) — https://stalepriors.com/runs/prisma--claude-opus-5--v1r-a--2026-09-01 prisma: places 6.7.0 (2025-04-29), cannot place 6.9.0 (2025-06-03) — https://stalepriors.com/runs/prisma--claude-opus-5--v1r-b--2026-09-01 prisma: places 6.7.0 (2025-04-29), cannot place 6.8.0 (2025-05-15) — https://stalepriors.com/runs/prisma--claude-opus-5--v2-a--2026-09-05 prisma: places 7.0.0 (2025-11-19), cannot place 7.1.0 (2025-12-03) — https://stalepriors.com/runs/prisma--claude-opus-5--v2-b--2026-09-05 prisma: places 6.8.0 (2025-05-15), cannot place 6.9.0 (2025-06-03) — https://stalepriors.com/runs/prisma--claude-opus-5--v3-a--2026-09-05 prisma: places 6.10.0 (2025-06-17), cannot place 6.11.0 (2025-07-01) — https://stalepriors.com/runs/prisma--claude-opus-5--v3-b--2026-09-05 prisma: places 7.0.0 (2025-11-19), cannot place 7.1.0 (2025-12-03) — https://stalepriors.com/runs/prisma--claude-opus-5--v3-c--2026-09-05 prisma: places 6.15.0 (2025-08-27), cannot place 6.16.0 (2025-09-10) — https://stalepriors.com/runs/prisma--claude-opus-5--v3-d--2026-09-05 prisma: places 7.0.0 (2025-11-19), cannot place 7.1.0 (2025-12-03) — https://stalepriors.com/runs/prisma--claude-opus-5--v3-e--2026-09-05 prisma: places 6.7.0 (2025-04-29), cannot place 6.8.0 (2025-05-15) — https://stalepriors.com/runs/prisma--claude-opus-5--v5-c--2026-09-06 prisma: places 7.0.0 (2025-11-19), cannot place 7.1.0 (2025-12-03) — https://stalepriors.com/runs/prisma--claude-opus-5--v5-d--2026-09-06 prisma: places 7.0.0 (2025-11-19), cannot place 7.1.0 (2025-12-03) — https://stalepriors.com/runs/prisma--claude-opus-5--v1--2026-08-31 tailwindcss: places 4.1.0 (2025-04-01), cannot place 4.2.0 (2026-02-18) — https://stalepriors.com/runs/tailwindcss--claude-opus-5--v2-a--2026-09-03 tailwindcss: places 4.1.0 (2025-04-01), cannot place 4.2.0 (2026-02-18) — https://stalepriors.com/runs/tailwindcss--claude-opus-5--v2-b--2026-09-03 tailwindcss: places 4.1.0 (2025-04-01), cannot place 4.2.0 (2026-02-18) — https://stalepriors.com/runs/tailwindcss--claude-opus-5--v3-c--2026-09-06 tailwindcss: places 4.1.0 (2025-04-01), cannot place 4.2.0 (2026-02-18) — https://stalepriors.com/runs/tailwindcss--claude-opus-5--v1--2026-08-31 valibot: places 1.0.0 (2025-03-19), cannot place 1.1.0 (2025-05-06) — https://stalepriors.com/runs/valibot--claude-opus-5--v1r-a--2026-09-01 valibot: places 1.0.0 (2025-03-19), cannot place 1.1.0 (2025-05-06) — https://stalepriors.com/runs/valibot--claude-opus-5--v1r-b--2026-09-01 valibot: places 1.0.0 (2025-03-19), cannot place 1.1.0 (2025-05-06) — https://stalepriors.com/runs/valibot--claude-opus-5--v2-a--2026-09-02 valibot: places 1.1.0 (2025-05-06), cannot place 1.2.0 (2025-11-24) — https://stalepriors.com/runs/valibot--claude-opus-5--v2-b--2026-09-02 valibot: places 1.1.0 (2025-05-06), cannot place 1.2.0 (2025-11-24) — https://stalepriors.com/runs/valibot--claude-opus-5--v3-c--2026-09-05 valibot: places 1.1.0 (2025-05-06), cannot place 1.2.0 (2025-11-24) — https://stalepriors.com/runs/valibot--claude-opus-5--v3-d--2026-09-05 valibot: places 1.1.0 (2025-05-06), cannot place 1.2.0 (2025-11-24) — https://stalepriors.com/runs/valibot--claude-opus-5--v4-c--2026-09-06 valibot: places 1.1.0 (2025-05-06), cannot place 1.2.0 (2025-11-24) — https://stalepriors.com/runs/valibot--claude-opus-5--v1--2026-09-01 zod: places 4.1.0 (2025-08-23), cannot place 4.2.0 (2025-12-15) — https://stalepriors.com/runs/zod--claude-opus-5--v2--2026-08-29 zod: places 4.1.0 (2025-08-23), cannot place 4.2.0 (2025-12-15) — https://stalepriors.com/runs/zod--claude-opus-5--v2r-a--2026-09-01 zod: places 4.1.0 (2025-08-23), cannot place 4.2.0 (2025-12-15) — https://stalepriors.com/runs/zod--claude-opus-5--v2r-b--2026-09-01 zod: places 4.1.0 (2025-08-23), cannot place 4.2.0 (2025-12-15) — https://stalepriors.com/runs/zod--claude-opus-5--v3-a--2026-09-02 zod: places 4.1.0 (2025-08-23), cannot place 4.2.0 (2025-12-15) — https://stalepriors.com/runs/zod--claude-opus-5--v3-b--2026-09-02 - Claude Sonnet 5 (states cutoff 2026-01): brackets from 7 libraries do NOT intersect — no single date explains this model. It describes langchain 1.0.0 (2025-10-17) but not next.js 15.1.0 (2024-12-10), published 311 days earlier. better-auth (7 draws), langchain (6 draws), next.js (4 draws) also produced brackets that exclude each other within one library; the verdict does not rest on that — no choice of a single draw per library admits any date either. better-auth: places 1.0.0 (2024-11-23), cannot place 1.1.0 (2024-12-20) — https://stalepriors.com/runs/better-auth--claude-sonnet-5--v2--2026-09-02 better-auth: places 1.2.0 (2025-03-01), cannot place 1.3.0 (2025-07-19) — https://stalepriors.com/runs/better-auth--claude-sonnet-5--v3-c--2026-09-03 better-auth: places 1.1.0 (2024-12-20), cannot place 1.2.0 (2025-03-01) — https://stalepriors.com/runs/better-auth--claude-sonnet-5--v4-a--2026-09-03 better-auth: places 1.0.0 (2024-11-23), cannot place 1.1.0 (2024-12-20) — https://stalepriors.com/runs/better-auth--claude-sonnet-5--v4-b--2026-09-03 better-auth: places 1.1.0 (2024-12-01), cannot place 1.2.0 (2025-03-01) — https://stalepriors.com/runs/better-auth--claude-sonnet-5--v5-c--2026-09-03 better-auth: places 1.0.0 (2024-11-23), cannot place 1.1.0 (2024-12-20) — https://stalepriors.com/runs/better-auth--claude-sonnet-5--v7-e--2026-09-06 better-auth: places 1.0.0 (2024-11-23), cannot place 1.1.0 (2024-12-20) — https://stalepriors.com/runs/better-auth--claude-sonnet-5--v1--2026-09-01 langchain: places 1.0.0 (2025-10-17), cannot place 1.1.0 (2025-11-24) — https://stalepriors.com/runs/langchain--claude-sonnet-5--v1r-a--2026-09-01 langchain: places 0.3.0 (2024-09-13), cannot place 1.0.0 (2025-10-17) — https://stalepriors.com/runs/langchain--claude-sonnet-5--v1r-b--2026-09-01 langchain: places 1.0.0 (2025-10-17), cannot place 1.1.0 (2025-11-24) — https://stalepriors.com/runs/langchain--claude-sonnet-5--v5-d--2026-09-06 langchain: places 0.3.0 (2024-09-13), cannot place 1.0.0 (2025-10-17) — https://stalepriors.com/runs/langchain--claude-sonnet-5--v6-c--2026-09-06 langchain: places 0.3.0 (2024-09-13), cannot place 1.0.0 (2025-10-17) — https://stalepriors.com/runs/langchain--claude-sonnet-5--v6-d--2026-09-06 langchain: places 0.3.0 (2024-09-13), cannot place 1.0.0 (2025-10-17) — https://stalepriors.com/runs/langchain--claude-sonnet-5--v1--2026-08-31 next.js: places 15.0.0 (2024-10-21), cannot place 15.1.0 (2024-12-10) — https://stalepriors.com/runs/next.js--claude-sonnet-5--v2-c--2026-09-02 next.js: places 15.3.0 (2025-04-09), cannot place 15.4.0 (2025-05-30) — https://stalepriors.com/runs/next.js--claude-sonnet-5--v3-a--2026-09-02 next.js: places 15.5.0 (2025-08-20), cannot place 16.0.0 (2025-10-22) — https://stalepriors.com/runs/next.js--claude-sonnet-5--v3-b--2026-09-02 next.js: places 15.0.0 (2024-10-21), cannot place 15.1.0 (2024-12-10) — https://stalepriors.com/runs/next.js--claude-sonnet-5--v1--2026-08-31 prisma: places 6.0.0 (2024-11-28), cannot place 6.1.0 (2024-12-17) — https://stalepriors.com/runs/prisma--claude-sonnet-5--v2-c--2026-09-05 prisma: places 6.0.0 (2024-11-28), cannot place 6.1.0 (2024-12-17) — https://stalepriors.com/runs/prisma--claude-sonnet-5--v5-a--2026-09-06 prisma: places 6.0.0 (2024-11-28), cannot place 6.1.0 (2024-12-17) — https://stalepriors.com/runs/prisma--claude-sonnet-5--v5-b--2026-09-06 prisma: places 6.0.0 (2024-11-28), cannot place 6.1.0 (2024-12-17) — https://stalepriors.com/runs/prisma--claude-sonnet-5--v1--2026-08-31 tailwindcss: places 4.1.0 (2025-04-01), cannot place 4.2.0 (2026-02-18) — https://stalepriors.com/runs/tailwindcss--claude-sonnet-5--v2-c--2026-09-03 tailwindcss: places 4.1.0 (2025-04-01), cannot place 4.2.0 (2026-02-18) — https://stalepriors.com/runs/tailwindcss--claude-sonnet-5--v3-d--2026-09-06 tailwindcss: places 4.1.0 (2025-04-01), cannot place 4.2.0 (2026-02-18) — https://stalepriors.com/runs/tailwindcss--claude-sonnet-5--v1--2026-08-31 valibot: places 1.0.0 (2025-03-19), cannot place 1.1.0 (2025-05-06) — https://stalepriors.com/runs/valibot--claude-sonnet-5--v2-c--2026-09-02 valibot: places 1.0.0 (2025-03-19), cannot place 1.1.0 (2025-05-06) — https://stalepriors.com/runs/valibot--claude-sonnet-5--v2-d--2026-09-02 valibot: places 1.0.0 (2025-03-19), cannot place 1.1.0 (2025-05-06) — https://stalepriors.com/runs/valibot--claude-sonnet-5--v3-e--2026-09-05 valibot: places 1.0.0 (2025-03-19), cannot place 1.1.0 (2025-05-06) — https://stalepriors.com/runs/valibot--claude-sonnet-5--v1--2026-09-01 zod: places 4.0.0 (2025-07-10), cannot place 4.1.0 (2025-08-23) — https://stalepriors.com/runs/zod--claude-sonnet-5--v2--2026-08-29 zod: places 4.0.0 (2025-07-10), cannot place 4.1.0 (2025-08-23) — https://stalepriors.com/runs/zod--claude-sonnet-5--v3-c--2026-09-02 zod: places 4.0.0 (2025-07-10), cannot place 4.1.0 (2025-08-23) — https://stalepriors.com/runs/zod--claude-sonnet-5--v4-a--2026-09-02 zod: places 4.0.0 (2025-07-10), cannot place 4.1.0 (2025-08-23) — https://stalepriors.com/runs/zod--claude-sonnet-5--v4-b--2026-09-02 ============================================================================== RUN better-auth--claude-fable-5-1--v6--2026-09-06 What Claude Fable 5.1 gets right about better-auth — battery v6, tested 2026-09-06 URL: https://stalepriors.com/runs/better-auth--claude-fable-5-1--v6--2026-09-06 JSON: https://stalepriors.com/data/better-auth/fable-5-1-v6.json Library: better-auth 1.7.3 (npm), verified 2026-09-06 Model: Claude Fable 5.1 (claude-fable-5-1), stated cutoff 2026-06, SELF-TEST Version attribution stops at: 1.3.0 (2025-07-19), lag ~11 months Oldest release it could not place: 1.4.0 (2025-11-22) Battery: better-auth/v6, 2 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: A boundary ladder, not a battery: two tasks, no code graded for correctness beyond the demonstration, nothing charged, `elicits_code: false`. It fills one of the five `?` cells BACKLOG 11k-b names. Claude Fable 5.1 describes better-auth through 1.3.0 (2025-07-19) and holds 1.4.0 onward as a bare number, eleven months below its stated June 2026 cutoff — the widest gap this subject has shown on any of the four libraries it has now been placed on. The poison rung was answered NAME rather than NO: the subject believes in a better-auth 0.9.0 that was never published, which does not void the ladder but is the arm's sharpest single line. --- non-finding [context] boundary: THE CELL THIS ARM WAS RUN TO FILL. Claude Fable 5.1 had never been drawn on better-auth, so `tools/charge-windows.mjs` printed `?` for every better-auth release against it. The boundary is 1.3.0 (2025-07-19) describable, 1.4.0 (2025-11-22) known as a number only — and all three of the arm's independent placements agree: the rung sort (1.2.0 and 1.3.0 DESCRIBE, 1.4.0 and 1.5.0 NAME, 1.6.0 and 1.7.0 NO), direct question (a), and direct question (c)'s explicit "first release I know only as a version number: 1.4.0". Consequence for the sweep: better-auth 1.4.0, 1.5.0 and 1.6.0 are all LIVE against this subject, and 1.6.0 (2026-04-06) carries no finding from any subject. --- non-finding [context] poison rung 0.9.0: THE POISON RUNG WAS ANSWERED `NAME`, NOT `NO`, AND THAT IS THE MOST INTERESTING LINE IN THE ARM. better-auth published no stable 0.9.0 — the 0.x minors run 0.8.0 (2024-11-08) straight to 1.0.0 (2024-11-23), fifteen days apart in a cadence of roughly one minor a week. The subject answered NAME and glossed it as "a late-2024 pre-1.0 release I believe existed but can't itemize". The ladder's stop rule voids an arm that claims to DESCRIBE the poison rung; NAME is inside the word's stated definition ("you know that version exists, OR BELIEVE IT DOES"), so the ladder holds and this is recorded rather than charged. It is still a belief in a release that never shipped, and it is a cheaper belief than it looks: a dense 0.x cadence with one gap in it is exactly where an interpolation costs nothing to make. --- non-finding [correct] TASK 1 — email/password auth instance and session guard: The demonstration task, and it demonstrates. `betterAuth({ database: new Pool(...), emailAndPassword: { enabled: true } })`, `toNodeHandler(auth)` mounted before `express.json()` with the raw-body reason given, `auth.api.getSession({ headers: fromNodeHeaders(req.headers) })` in the guard, and the Express 5 `/api/auth/*splat` route-syntax caveat volunteered unprompted. This is the 1.x API and it is what the section exists for: the subject demonstrates before it is asked to commit to a boundary (prompts/boundary-ladder.md). --- non-finding [correct] release dating below the boundary: Four consecutive minors dated correctly inside direct question (c): 1.0.0 ~Nov 2024 (2024-11-23), 1.1.0 ~Dec 2024 (2024-12-20), 1.2.0 ~Mar 2025 (2025-03-01), 1.3.0 ~Jul 2025 (2025-07-19). Below its own boundary this subject's attribution is accurate to the month on every rung it claimed, which is what licenses reading the boundary it then stated. ============================================================================== RUN better-auth--claude-fable-5-1--v7-c--2026-09-06 What Claude Fable 5.1 gets wrong about better-auth — battery v7-c, tested 2026-09-06 URL: https://stalepriors.com/runs/better-auth--claude-fable-5-1--v7-c--2026-09-06 JSON: https://stalepriors.com/data/better-auth/fable-5-1-v7-c.json Library: better-auth 1.7.3 (npm), verified 2026-09-06 Model: Claude Fable 5.1 (claude-fable-5-1), stated cutoff 2026-06, SELF-TEST Version attribution stops at: 1.4.0 (2025-11-22), lag ~7 months Oldest release it could not place: 1.5.0 (2026-03-01) Battery: better-auth/v7-c, 7 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The second test arm, and the one that charged the surface the battery was built for. It answered task 1 "yes" and task 5 `updatedAt` - a coherent, confidently argued account of better-auth 1.5.0 offered as the current behaviour - and denied both 1.6.0 options exist. Three findings: F1 the freshness anchor (charged once across both tasks, as the pre-registration required), F2 `resendStrategy`, F3 `twoFactorPage`. F2 is the sharpest of the three because the workaround it shipped sets `storeOTP: 'hashed'` and then caches the plaintext code in Redis to make reuse possible, which is precisely the configuration in which the real option refuses to reuse. Its stated cutoff is June 2026, two months after the release it missed. --- F1 [S2 silently-wrong] Predicts the freshness check passes for a 30-hour-old session that was used two minutes ago, and names `updatedAt` as the anchor — the pre-1.6.0 semantics, held consistently across two tasks API: session.freshAge (measured from session.createdAt) Changed in better-auth 1.6.0 (2026-04-06), kind: behavior-changed Chargeability: 1.6.0 shipped 2026-04-06; this draw states a June 2026 cutoff, which is after it and not the same month, so the fairness rule and the same-month bar both clear. This is a pre-registered test arm and tasks 1 and 5 are pre-registered probes on this surface. **One finding, two artefacts.** The pre-registration graded the verdict (task 1) and the mechanism (task 5) independently and this draw got both wrong in the same direction; that is one belief measured twice, not two findings, and charging it twice would inflate the count. Model belief: Task 1, verdict on its own line: "yes". Task 5, four tasks later: "(a) `updatedAt` — with a fallback to `createdAt` if `updatedAt` is absent on the row." That is a verbatim description of the 1.5.0 implementation, which read `new Date(session.session.updatedAt || session.session.createdAt)`. At task 3 it restated the same belief as the reason its own answer to task 1 was "yes": "As best I remember the middleware, it measures from `updatedAt` and only falls back to `createdAt` when `updatedAt` is missing — which is why Task 1 passes. That is the product team's semantics, not the security team's, and it's baked in." It also flagged, correctly, that the vendor's documentation and the implementation had historically disagreed, and told the reader to check `freshSessionMiddleware` in the installed package - which would have corrected it. Wrong: // The advice that follows from the belief: keep the built-in check as a // 'last used' gate, and add your own createdAt gate on top for the // endpoints security cares about. export const auth = betterAuth({ session: { // Built-in check: measured from updatedAt (falls back to createdAt). freshAge: FRESH_AGE, expiresIn: 60 * 60 * 24 * 7, updateAge: 60 * 60 * 24, }, hooks: { before: createAuthMiddleware(async (ctx) => { /* createdAt gate */ }) }, }) Correct: // 1.6.0 and later the window runs from createdAt. Executed against three // installed releases with createdAt = now-30h and updatedAt = now-2min: // 1.5.0 -> the freshness gate PASSES (the call fails later, in the handler) // 1.6.0 -> FORBIDDEN / SESSION_NOT_FRESH // 1.7.3 -> FORBIDDEN / SESSION_NOT_FRESH // There is no option to move the anchor back. If you want last-use // semantics, disable the built-in check and write your own gate: export const auth = betterAuth({ session: { freshAge: 0 } }) Impact: Everything compiles and runs, which is what makes it S2 and what makes it hard to catch: the reader is given a correct prediction about better-auth 1.5.0 and told it is the current behaviour, with nothing in their own code to change and no deprecation warning to trip over. Two concrete consequences. **(1) The security posture is inverted.** The draw tells a team that the built-in gate is "the product team's semantics" - lenient, effectively never firing for an active user - and has them build a second, stricter `createdAt` gate in a `hooks.before` middleware with a hand-maintained list of sensitive paths, which the draw itself called "the honest weak spot". From 1.6.0 the built-in gate already is the `createdAt` gate; the custom middleware duplicates it, and the path list will silently miss whichever endpoints the framework adds to `freshSessionMiddleware` later. **(2) The upgrade surprise is not predicted.** A team upgrading 1.5.x -> 1.6.x on this advice expects nothing to change for continuously active users. Executed against the installed packages: a session with `createdAt` 30 hours ago and `updatedAt` 2 minutes ago passes the gate at 1.5.0 and is rejected with `FORBIDDEN / SESSION_NOT_FRESH` at 1.6.0 and 1.7.3. Source: https://registry.npmjs.org/better-auth/-/better-auth-1.6.0.tgz (2026-04-06) — "const createdAt = new Date(session.session.createdAt).getTime(); const freshAge = ctx.context.sessionConfig.freshAge * 1e3; if (Date.now() - createdAt >= freshAge) throw APIError.from("FORBIDDEN", BASE_ERROR_CODES.SESSION_NOT_FRESH);" Source: https://registry.npmjs.org/better-auth/-/better-auth-1.5.0.tgz (2026-03-01) — "const lastUpdated = new Date(session.session.updatedAt || session.session.createdAt).getTime();" Source: https://github.com/better-auth/better-auth/releases/tag/v1.6.0 (2026-04-06) — "Aligned `freshAge` calculation with session creation time instead of update time (#8762)" --- F2 [S2 silently-wrong] Denies the email-OTP plugin has a resend-reuse option and rebuilds it out of `generateOTP` plus a Redis cache — under `storeOTP: 'hashed'`, which is the one configuration in which the real option refuses to reuse API: emailOTP({ resendStrategy }) Changed in better-auth 1.6.0 (2026-04-06), kind: added Chargeability: Same licence as F1: stated cutoff June 2026, target release 2026-04-06, different months. Model belief: Verdict on its own line: "no". Then the explicit absence claim the additive-API rule requires: "I don't know of any option on `emailOTP()` that says 'reuse the outstanding code on resend.'" It then listed the plugin's options from memory - a list correct for 1.5.0 and missing only the one asked about - and described the mechanism correctly: "Every call to `sendVerificationOTP` replaces the stored verification row... with a freshly generated code, which is exactly the behaviour your support team is hitting." Wrong: // shipped as the answer. Note the third line of options. emailOTP({ otpLength: 6, expiresIn: OTP_TTL, allowedAttempts: 5, storeOTP: 'hashed', async generateOTP({ email, type }) { const existing = await redis.get(key(type, email)) if (existing) return existing // plugin re-stores it with a fresh expiresIn return String(Math.floor(100000 + Math.random() * 900000)) }, async sendVerificationOTP({ email, otp, type }) { await redis.set(key(type, email), otp, 'EX', OTP_TTL) await sendEmail({ to: email, text: `Your code is ${otp}.` }) }, }) Correct: // 1.6.0 and later: one option, and it already knows the constraint the // hand-rolled cache does not. emailOTP({ otpLength: 6, expiresIn: 300, resendStrategy: 'reuse', // default is 'rotate' storeOTP: 'plain', // reuse needs a recoverable code; 'hashed' falls back to 'rotate' async sendVerificationOTP({ email, otp }) { await sendMail(email, otp) }, }) Impact: The workaround runs, so S2 - but this draw's version is the one that shows why the built-in option is worth having. It sets `storeOTP: 'hashed'` and then keeps the plaintext code in Redis so it can hand it back, which puts the secret in a second store the security control was chosen to keep it out of. `resendStrategy: 'reuse'` does not have that failure mode available to it: the option's own contract is that reuse works only when the stored code is recoverable and it **falls back to `"rotate"` when the OTP is hashed**, so the library refuses to do quietly what this configuration does explicitly. A reader following this answer ends up with hashed storage in the database, plaintext in Redis, and the belief that they have hardened the flow. The draw also carried three honest caveats it should not have needed - whether `generateOTP` may be async, whether the code is cleared on successful verification, and that the attempt counter resets on every resend - all of which are questions about a re-implementation of a shipped option. Source: https://registry.npmjs.org/better-auth/-/better-auth-1.6.0.tgz (2026-04-06) — "- `"reuse"`: Resends the same OTP and extends its expiry. * Only works when the OTP is recoverable (plain, encrypted, or custom encrypt/decrypt). * Falls back to `"rotate"` when OTP is hashed." Source: https://registry.npmjs.org/better-auth/-/better-auth-1.5.0.tgz (2026-03-01) — "sendVerificationOTP, otpLength, expiresIn, generateOTP, sendVerificationOnSignUp, disableSignUp, allowedAttempts, storeOTP, changeEmail, overrideDefaultEmailVerification, rateLimit" --- F3 [S3 deprecated] Denies the two-factor client plugin takes a page option, names `twoFactorPage` correctly, and dates it to the pre-1.0 releases as something since removed API: twoFactorClient({ twoFactorPage }) Changed in better-auth 1.6.0 (2026-04-06), kind: added Chargeability: Same licence as F1 and F2. Severity capped at S3 in the pre-registration because `onTwoFactorRedirect` still exists, so the shipped code works. Model belief: Verdict on its own line: "no". Then: "the current client plugin takes a callback, `onTwoFactorRedirect()`, not a path. I believe there was a `twoFactorPage: \"/two-factor\"` string option on the client plugin in the pre-1.0 (0.x) days, and that's probably where the memory of 'give it a path once' comes from; I'm fairly, not fully, sure it's gone from 1.x." The name is right and the history is backwards. **This is the second model family to produce that inversion independently in this battery** - `v7-a` (Claude Opus 5) wrote "I have a real memory of a string-path option in early better-auth two-factor docs — I believe it was called `twoFactorPage` — but I think it was replaced by the callback" - and `v7-d` made it three of six draws. Wrong: // shipped as the answer, with the string option explicitly placed in the past. export const authClient = createAuthClient({ plugins: [ twoFactorClient({ onTwoFactorRedirect() { window.location.href = '/auth/two-factor' }, }), ], }) Correct: // 1.6.0 and later — the string option the task asked for: export const authClient = createAuthClient({ plugins: [twoFactorClient({ twoFactorPage: '/auth/two-factor' })], }) // The callback is still supported and since 1.6.0 is told which factors the // user has, which is the reason to keep using it in a router-driven app: twoFactorClient({ onTwoFactorRedirect({ twoFactorMethods }) { router.push(twoFactorMethods?.includes('totp') ? '/auth/totp' : '/auth/otp') }, }) Impact: S3 by design: the code works, and a reader loses only the one-line setup they asked for. The cost is in the story attached to it. Told that `twoFactorPage` is a name from the 0.x era that 1.x dropped, a reader will not try it and will read the current documentation as stale if they meet it. The draw's fallback - registering a module-level navigator function and calling `setNavigator(navigate)` from a top-level component - is fifteen lines of indirection standing in for a string. Source: https://registry.npmjs.org/better-auth/-/better-auth-1.6.0.tgz (2026-04-06) — "twoFactorPage?: string;" Source: https://registry.npmjs.org/better-auth/-/better-auth-1.5.0.tgz (2026-03-01) — "onTwoFactorRedirect?: () => void | Promise;" --- non-finding [correct] session.freshAge anchor option (does not exist): Task 3, the pre-registered control. "no", correctly - there is no option choosing the freshness anchor at any release. The reason it gave is the stale belief from F1 rather than knowledge of the option surface ("there is a single knob, `session.freshAge`... There is no option selecting which timestamp it's measured from"), but the graded answer is the right one and the control holds in this arm. --- non-finding [correct] customSession: Task 4, the floor probe (1.0.0 custom session). `customSession` on the server, `customSessionClient()` on the client, with the typing constraint stated correctly - pass the options object into `customSession(fn, options)` so plugin-widened `user`/`session` types reach the callback. Passed. --- non-finding [correct] session.freshAge: Task 5(b): `session: { freshAge: 0 }`, correct, and the only half of task 5 this draw got right. All six draws in the battery except the below-floor guessing control gave this answer, which is unsurprising - `0` as the off switch predates the window and is not what this battery probes. --- non-finding [context] stateless sessions: Task 7, the attribution anchor. Unlike both Opus draws this one declined rather than denying: "I can't name one with confidence, and I'd rather say that than invent it... If a true stateless-session mode shipped, it would be in a release after those (1.5 or later, 2026), and I don't know its contents. Treat this as 'unknown,' not 'doesn't exist.'" It then listed `cookieCache`, `secondaryStorage` with `storeSessionInDatabase: false`, and the `jwt` plugin as the near neighbours, and attributed a cookie-cache `strategy` option to 1.4 - the release that actually carries stateless session management. So it placed the right release for an adjacent feature while marking the feature itself unknown. Belief data, never scored. An abstention is not a denial (JOURNAL/046). --- non-finding [context] knowledge boundary: Boundary spread inside one subject on one library, worth recording because `better-auth/v6` measured this cell six hours earlier. The ladder placed Claude Fable 5.1's better-auth boundary at **1.3.0**; this draw describes **1.4.0** in detail ("several plugins split into their own packages... cookie-cache `strategy` options, adapter-level joins work") and names 1.5 as the first release it knows only as a number. Its twin `v7-d` agrees with the ladder at 1.3.0. So the subject's spread on this library is 1.3.0 / 1.3.0 / 1.4.0 across three draws, and this arm is the high one. Nothing charged here depends on it: 1.6.0 is above every reading. ============================================================================== RUN better-auth--claude-fable-5-1--v7-d--2026-09-06 What Claude Fable 5.1 gets right about better-auth — battery v7-d, tested 2026-09-06 URL: https://stalepriors.com/runs/better-auth--claude-fable-5-1--v7-d--2026-09-06 JSON: https://stalepriors.com/data/better-auth/fable-5-1-v7-d.json Library: better-auth 1.7.3 (npm), verified 2026-09-06 Model: Claude Fable 5.1 (claude-fable-5-1), stated cutoff 2026-06, SELF-TEST Version attribution stops at: 1.3.0 (2025-07-19), lag ~11 months Oldest release it could not place: 1.4.0 (2025-11-22) Battery: better-auth/v7-d, 7 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The blind twin, and it agreed with `v7-c` on every graded answer: task 1 "yes", task 5 `updatedAt`, both 1.6.0 options denied, the control held, the floor passed. That agreement is what makes the Fable 5.1 freshness belief a measurement rather than a coin - and it is the direct contrast with the Opus 5 pair, which split on the same question from the same prompt. Two details are its own. It required `storeOTP: 'plain'` for its resend workaround, deriving the exact constraint `resendStrategy` encodes without knowing the option exists. And it stated the `twoFactorPage` history inversion flatly rather than as a hedged memory. It charges nothing, by the duplicated-arm rule. --- non-finding [miss] session.freshAge (measured from session.createdAt): Task 1: "yes", reproducing its twin exactly, and consistent with its own task 5(a) answer of `updatedAt`. At task 3 it stated the belief outright: "the timestamp it is compared against is fixed in the framework's `freshSessionMiddleware` — it uses `updatedAt`, i.e. the product team's semantics. Because `updateAge` refreshes `updatedAt` on activity, an active user is effectively always 'fresh' under the default." That is an accurate description of better-auth up to 1.5.0. Both Fable 5.1 draws agree here, which is what turns the belief into a measurement; the two Opus 5 draws split on the same question. --- non-finding [miss] emailOTP({ resendStrategy }): Task 2: "no", with the absence claimed explicitly ("Nothing among them says 'reuse the outstanding code on resend'") after listing the 1.5.0 option set from memory. Shipped the same `generateOTP` reuse cache as its twin, but read the code back out of better-auth's own verification row through `internalAdapter.findVerificationValue` rather than out of Redis - and, unlike its twin, correctly required `storeOTP: 'plain'` for it to be possible at all, noting that "'hashed' makes this impossible". That is the constraint `resendStrategy` encodes, derived correctly by a draw that did not know the option exists. --- non-finding [miss] twoFactorClient({ twoFactorPage }): Task 6: "no", with the same inverted history as `v7-c` and `v7-a`: "Early 0.x releases accepted `twoFactorClient({ twoFactorPage: '/two-factor' })`; that string option was replaced by the `onTwoFactorRedirect` callback and is what the current release supports." This draw states it flatly rather than hedging it. Three of the battery's six draws name `twoFactorPage` correctly and place it on the wrong side of its own introduction. --- non-finding [correct] session.freshAge anchor option (does not exist): Task 3, the control: "no". Correct, and it went further than any other arm in shipping the alternative - a `MEASURE_FROM` constant driving both `freshAge` and a `hooks.before` middleware, so that the choice reads as a setting in the reader's own config. It named the weak spot itself: "The path list is the honest weak spot — you have to keep it in sync with which endpoints the framework guards with `freshSessionMiddleware`." --- non-finding [correct] customSession: Task 4, the floor probe. `customSession` with the client plugin for typing, plus the `cookieCache` interaction and the `session.additionalFields` alternative for a stored rather than computed value. Passed. --- non-finding [context] stateless sessions: Task 7: "I don't know of a release that does this, and I'd be suspicious of any answer that names one." Then the same three near neighbours the other draws listed, and a note that later 1.x releases extended cookie caching "I have a vague memory of a strategy/refresh option for the cache in the 1.4–1.5 era" - the right window for the wrong feature. Stateless session management is 1.4.0. Belief data, never scored. ============================================================================== RUN better-auth--claude-fable-5--v2-a--2026-09-02 What Claude Fable 5 gets wrong about better-auth — battery v2-a, tested 2026-09-02 URL: https://stalepriors.com/runs/better-auth--claude-fable-5--v2-a--2026-09-02 JSON: https://stalepriors.com/data/better-auth/fable-5-v2-a.json Library: better-auth 1.7.2 (npm), verified 2026-09-02 Model: Claude Fable 5 (claude-fable-5), stated cutoff 2026-01 Version attribution stops at: 1.3.0 (2025-07-19), lag ~6 months Oldest release it could not place: 1.4.0 (2025-11-22) Battery: better-auth/v2-a, 5 tasks, tool uses during test: 0 Tested: 2026-09-02 Summary: The cleanest instance of the effect this battery was built to look for. Claude Fable 5 wired both of the plugins better-auth shipped in the patch release 1.3.8 — device authorization and last-login-method — correctly and from real knowledge, then attributed both to the minor 1.3.0, alongside three features that genuinely are 1.3.0. It passed the battery's internal control by placing SAML SSO at 1.3.0 correctly, which is what makes the misplacement readable as a collapse rather than as noise. It also volunteered the effect while committing it: "I may be attributing one or two items to it that actually landed in a nearby 1.3.x patch." One charged finding, S2: it stated plainly that better-auth has no store-free session mode, two months after its own stated cutoff should have covered the 1.4.0 release that added one. --- F1 [S2 silently-wrong] States that better-auth offers no stateless, store-free session mode, and names cookieCache, Redis and the jwt plugin as the only options API: stateless / database-less sessions Changed in better-auth 1.4.0 (2025-11-22), kind: added Chargeability: 1.4.0 (2025-11-22) precedes this subject's stated cutoff of 2026-01 by two months. Model belief: "Here I want to be plain rather than oversell: better-auth does not, to my knowledge, offer a fully stateless 'the cookie is the session' mode with no backing store at all. ... So: if the requirement is literally 'cookie is the only place the session lives, zero server-side store,' I don't believe the library supports that as a first-class mode, and I'd say so to the team rather than pretend cookieCache is it." Wrong: // all three options offered keep a server-side store session: { cookieCache: { enabled: true, maxAge: 300 } } export const auth = betterAuth({ secondaryStorage: redisStore, session: { storeSessionInDatabase: false }, }) Correct: // stateless since 1.4.0: omit BOTH `database` and `secondaryStorage` export const auth = betterAuth({ emailAndPassword: { enabled: true }, }) Impact: The reader is told to shrink the cookie or stand up Redis, and is explicitly advised that the thing they asked for does not exist. It has existed since 2025-11-22. The subject's care in flagging the denial as deliberate rather than hedged makes the advice more likely to be acted on, not less. Scope note: Verified at both ends per fact LF1. Not verified: that the feature is unchanged between 1.4.0 and 1.7.2. Source: https://github.com/better-auth/better-auth/releases/tag/v1.4.0 (2025-11-22) — "Stateless session management" Source: https://registry.npmjs.org/better-auth/-/better-auth-1.7.2.tgz (2026-08-26) — "the only place the session lives and therefore the authority itself (`false`, for stateless / DB-less deployments)" --- non-finding [context] deviceAuthorization() / lastLoginMethod(): THE MEASUREMENT, and this draw is the clearest instance of it in the battery. The internal control passed: SAML enterprise SSO placed at "1.3.0", correct (fact LF4). Having shown it can attribute a genuine minor, its patch answers can be read — and both patch-shipped plugins were placed at the same minor. Device authorization: "introduced in 1.3.0 — estimate; I'm confident it's the 1.3 line, less confident it was the .0 rather than an early 1.3.x patch." Last-login-method: "1.3.0 — estimate ... I strongly associate lastLoginMethod with the 1.3 announcement." Both shipped in 1.3.8. Its list in (c) makes the shape unmistakable: it names 1.3 as containing "SAML support in the SSO plugin, deviceAuthorization plugin, lastLoginMethod plugin, Sign in with Ethereum (SIWE), multi-team support" — of which SAML, SIWE and multi-team really are 1.3.0 and the two plugins are 1.3.8, absorbed into the minor beside them. --- non-finding [context]: The subject named the effect itself, unprompted, while committing it. Its own caveat on the 1.3 line reads: "I can describe this release's contents with reasonable confidence, though I may be attributing one or two items to it that actually landed in a nearby 1.3.x patch." It was attributing exactly two such items, and it could not tell which two. That is the collapse described from the inside: the model knows its attribution has minor granularity and cannot resolve below it. --- non-finding [correct] deviceAuthorization() / lastLoginMethod(): Tasks 1 and 2 passed on capability: `deviceAuthorization()` and `deviceAuthorizationClient()` with the RFC 8628 code/poll/approve shape, and `lastLoginMethod()` with `lastLoginMethodClient()`, `getLastUsedLoginMethod()`, `isLastUsedLoginMethod()`, the cookie default and the per-browser caveat. Both are 1.3.8 surface (facts LF5, LF6), so the knowledge is present and correct and only the version attached to it is wrong. --- non-finding [imprecision] additional user fields in the sign-in response: Task 4 falsified prediction P3's second half. The handler reads `const plan = data.user.plan` off the sign-in response — correct since 1.4.2 (2025-11-25) — with the prose hedge "additional fields are definitely returned on getSession. My recollection is the signIn.email response's user also carries them, but I'm not 100% certain across versions." Right code, stated as uncertain. --- non-finding [correct] customSession(): Task 5, the floor probe, passed: `customSession()` with `customSessionClient()` and the unprompted note that whatever it returns is serialized into the cookie cache, which loops back to the cookie-size problem in task 3. The 1.0.0 floor is confirmed, so this run's boundary reading is a measurement. ============================================================================== RUN better-auth--claude-fable-5--v2-b--2026-09-02 What Claude Fable 5 gets right about better-auth — battery v2-b, tested 2026-09-02 URL: https://stalepriors.com/runs/better-auth--claude-fable-5--v2-b--2026-09-02 JSON: https://stalepriors.com/data/better-auth/fable-5-v2-b.json Library: better-auth 1.7.2 (npm), verified 2026-09-02 Model: Claude Fable 5 (claude-fable-5), stated cutoff 2026-01 Version attribution stops at: 1.3.0 (2025-07-19), lag ~6 months Oldest release it could not place: 1.4.0 (2025-11-22) Battery: better-auth/v2-b, 5 tasks, tool uses during test: 0 Tested: 2026-09-02 Summary: The concurrent blind twin of `v2-a`, and the draw that states the effect most starkly. It placed its own knowledge boundary inside better-auth's 1.3 patch line — "the 1.3.x patches beyond roughly 1.3.2-1.3.4" are version numbers with no content attached — and then, in the next answer, attributed two plugins that shipped in 1.3.8 to 1.3.0, calling it "a firm estimate" and explicitly ruling out a patch. It agreed with its twin on every task and on all four attributions. It charges nothing, per the rule that the duplicated arm does not charge; its denial of database-less sessions is recorded as a chargeable miss carried by its twin. --- non-finding [context] deviceAuthorization() / lastLoginMethod(): THE MEASUREMENT, and the sharpest single answer the battery produced. This draw declared a boundary INSIDE the 1.3 patch line and then reached across it. In (c) it named its own contentless region as "the 1.3.x patches beyond roughly 1.3.2-1.3.4". In (d), asked where device authorization was introduced, it answered "1.3.0 (~July 2025). Fairly confident it was a 1.3.0 headline feature, not a patch. Estimate, but a firm one" — and gave last-login-method "same release, same confidence". Both shipped in 1.3.8: a release inside the very range it had just described as version numbers with no content attached. It did not merely fail to reach the patch; it explicitly ruled the patch out. The internal control passed — SAML placed at 1.3.0, correct (fact LF4) — so this is a collapse and not noise. --- non-finding [miss] stateless / database-less sessions: Task 3: "better-auth does not have a fully database-free mode", followed by cookieCache, secondaryStorage and the jwt/bearer plugins, and closing "I don't believe there's a supported stateless-only session mode; if that's a hard requirement I'd say so plainly to the team rather than fight the library." Fact LF1 records that 1.4.0 (2025-11-22) added exactly that, two months before this subject's stated 2026-01 cutoff. Chargeable, and charged on the `v2-a` draw. --- non-finding [correct] deviceAuthorization() / lastLoginMethod() / customSession(): Tasks 1, 2 and 5 passed and matched the twin draw closely: `deviceAuthorization()` with the RFC 8628 poll loop, `lastLoginMethod()` with the non-httpOnly cookie and `storeInDatabase` option, and `customSession()` for the floor probe with the note to keep it last in the plugins array. Floor confirmed, so the boundary reading is a measurement. --- non-finding [imprecision] additional user fields in the sign-in response: Task 4, against prediction P3: the handler reads `data.user.plan` off the sign-in response — correct since 1.4.2 — then adds a defensive `getSession()` fallback behind a null check, with the comment "in some versions the sign-in response user has been thinner than the session user" and the note "I remember issue traffic about whether the signIn.email response user carries additional fields in all versions." Four of four test-arm draws produced this same pattern: correct code, disbelieved in prose. ============================================================================== RUN better-auth--claude-fable-5--v3-d--2026-09-03 What Claude Fable 5 gets right about better-auth — battery v3-d, tested 2026-09-03 URL: https://stalepriors.com/runs/better-auth--claude-fable-5--v3-d--2026-09-03 JSON: https://stalepriors.com/data/better-auth/fable-5-v3-d.json Library: better-auth 1.7.2 (npm), verified 2026-09-03 Model: Claude Fable 5 (claude-fable-5), stated cutoff 2026-01 Version attribution stops at: 1.3.0 (2025-07-19), lag ~18 months Oldest release it could not place: 1.4.0 (2025-11-22) Battery: better-auth/v3-d, 4 tasks, tool uses during test: 0 Tested: 2026-09-03 Summary: The second below-floor control, and the draw that settled the battery's other half. It failed the `baseURL` probe like every other arm, which is what licenses reading the two Opus failures as beliefs rather than as noise. But on verification hashing it answered yes and named `storeOTP` and `storeToken` with their correct value sets - from eighteen months below the target release - which is what proved the Index's own newly written fact LF9 had claimed too much. It also gave the most accurate version attribution of the four draws, placing SAML SSO exactly at 1.3.0. A control arm that corrects the operator is a better outcome than one that merely confirms a floor. --- non-finding [miss] baseURL as a dynamic multi-host config: Task 1, the second control result: "(i) No. `baseURL` is typed as a plain optional string. There is no function form, no array form." It then proposed dropping `baseURL` entirely so the library infers the origin from the request, gated by `trustedOrigins` - a different workaround from the two Opus draws and, notably, the closest of the four to the shape of the real feature without ever reaching it. 1.5.0 is two months above this subject's stated cutoff. Prediction P2 confirmed on this arm too: neither Opus denial can be an artefact of an unguessable name. --- non-finding [correct] verification.storeIdentifier: THE CONTROL THAT ANSWERED THE PROBE THE TEST ARM WAS SUPPOSED TO FAIL. On task 3 this below-floor draw answered "(i) Yes - but per plugin, not as one global switch on the `verification` table. The email-OTP plugin has a `storeOTP` option and the magic-link plugin grew an equivalent `storeToken`; both accept \"plain\" (default), \"hashed\", and (for OTP at least) \"encrypted\", plus a custom-hasher form." Every clause of that is correct against the published 1.3.0 declarations, including that `encrypted` exists for OTP and not for magic-link. Its (iii) answer on deterministic hashing is correct. Per JOURNAL/030 the outcome of this probe is therefore DERIVABLE-at-1.3.0 for the plugin half - which is not a disqualification of anything, because the plugin half is genuinely old knowledge; it is the evidence that made the Index correct its own fact LF9 the same session. --- non-finding [correct] session token hashing at rest: Task 2, the sibling control: "(i) No. To my knowledge there is no config that hashes the `session.token` column before write", plus the unprompted and correct observation that a `databaseHooks.session.create` hook cannot fix it because the library's own lookup would then miss. Nothing invented. P3 holds on all four arms. --- non-finding [correct] customSession(): Task 4, the floor probe, passed: `customSession()` plus `customSessionClient()`, with the correct note that it runs on every `getSession` and that a stored rather than computed field belongs in `additionalFields`. Floor confirmed on all four arms. --- non-finding [context]: The most accurate belief data of the four draws. It placed SAML SSO at "1.3, ~July 2025" - fact LF4 records 1.3.0, 2025-07-19, exactly right - and placed the verification-hashing options in "the 1.2.x line, roughly mid-2025 ... around 1.2.7", which is one minor early against the real 1.3.0 but is the only draw to give a number at all. Its (d)(iii) denial of database-less operation matches every other draw against fact LF1 (1.4.0), which is two months above its stated cutoff and therefore not chargeable here. ============================================================================== RUN better-auth--claude-fable-5--v4-e--2026-09-03 What Claude Fable 5 gets right about better-auth — battery v4-e, tested 2026-09-03 URL: https://stalepriors.com/runs/better-auth--claude-fable-5--v4-e--2026-09-03 JSON: https://stalepriors.com/data/better-auth/fable-5-v4-e.json Library: better-auth 1.7.2 (npm), verified 2026-09-03 Model: Claude Fable 5 (claude-fable-5), stated cutoff 2026-01 Version attribution stops at: 1.3.0 (2025-07-19), lag ~18 months Oldest release it could not place: 1.4.0 (2025-11-22) Battery: better-auth/v4-e, 3 tasks, tool uses during test: 0 Tested: 2026-09-03 Summary: Claude Fable 5's charging arm passed the target probe cleanly and denied the sibling correctly, which is the discrimination pattern that reads as recall rather than as running the naming scheme. It named `storeToken` and `storeOTP`, gave the right defaults, pointed the reader at the real `PhoneNumberOptions` type name for the plugin that lacks the option, and reproduced better-auth's internal `code:attemptCount` verification-value format from memory — a string this session confirmed by execution. Nothing is charged. Its one clean miss is attribution: it dated the options to about 1.2.7 while, in the same answer, correctly dating SAML SSO to 1.3.0 — the release the options actually shipped in. --- non-finding [correct] magicLink storeToken / emailOTP storeOTP: Task 2, the target probe: "Yes, both. The magic-link plugin takes `storeToken`, and the email-OTP plugin takes `storeOTP`. Both accept `"plain"` (the default), `"hashed"`, `"encrypted"`, and a custom variant where you supply your own hash or encrypt/decrypt functions." Correct on the names and the defaults. One over-generalisation inside a correct answer: `"encrypted"` and the `{ encrypt, decrypt }` pair are in `storeOTP`'s union and not in `storeToken`'s, which is `"plain" | "hashed" | { type: "custom-hasher", hash }`. The draw flagged its own uncertainty about the discriminant names and told the reader to check the types. --- non-finding [correct] phoneNumber storeOTP (does not exist): Task 1, the same-scheme sibling control: correct denial, no invention, and it named the exact type to check — "check the installed version's option types (`PhoneNumberOptions`) before taking my 'no' as final". `PhoneNumberOptions` is the real name of that interface at 1.7.2 and it contains no storage option. It also declined the `databaseHooks` workaround for the right reason: "the plugin does the plaintext comparison internally, so a `databaseHooks.verification.create` hook that hashes the value would just break `verifyPhoneNumber`." --- non-finding [correct] verification value format (code:attemptCount): Recalled the library's internal storage format for a phone OTP unprompted: "better-auth actually stores it as `code:attemptCount` internally in the versions I know". Executed against an installed 1.3.0, the email-OTP verification row holds `797478:0` — the code, a colon, and the attempt counter. The format is exact. --- non-finding [miss] magicLink storeToken / emailOTP storeOTP: Version attribution, direct question (d)(i)/(ii): placed the per-plugin hashing options in the 1.2.x line — "Shipped in a 1.2.x patch, ~April-May 2025 — my best estimate is around 1.2.7." They shipped in **1.3.0** (2025-07-19). `storeOTP` and `storeToken` appear nowhere in the published `dist` of `better-auth@1.2.7` or `better-auth@1.2.12` — 1.2.12 being the last stable 1.2.x — and appear in five files of 1.3.0's. The subject demonstrated it holds the capability and then dated it one minor low, which is the case HARNESS.md's attribution rule says is readable. In the same answer this draw dated SAML support in the SSO plugin to 1.3.0, ~July 2025, which is correct and is the same release the hashing options shipped in. One transcript, one release, two features, one dated right and one dated a minor low. --- non-finding [miss] twoFactor otpOptions.storeOTP: Direct question (d)(iii) denied that the two-factor plugin has an at-rest hashing option for its OTP: "I do not know of this option existing... the sent OTP fallback in the two-factor plugin I believe is stored raw, with no hashing option in versions I can describe." It does. `twoFactor({ otpOptions: { storeOTP } })` is present in `better-auth@1.3.0` and type-checks under `tsc --strict` at 1.3.0, 1.5.0 and 1.7.2 — it is the third of the four plugins that gained the option in that release. This is a real gap inside the fairness window and it is not charged, because the battery's pre-registration states that the direct questions are belief data and are never scored as findings. --- non-finding [correct] customSession: Floor probe (task 3) passed: named the `customSession` plugin, spread `user` and `session` back out of the callback, and added the companion `customSessionClient` on the client for type inference. The run is readable. ============================================================================== RUN better-auth--claude-fable-5--v4-f--2026-09-03 What Claude Fable 5 gets right about better-auth — battery v4-f, tested 2026-09-03 URL: https://stalepriors.com/runs/better-auth--claude-fable-5--v4-f--2026-09-03 JSON: https://stalepriors.com/data/better-auth/fable-5-v4-f.json Library: better-auth 1.7.2 (npm), verified 2026-09-03 Model: Claude Fable 5 (claude-fable-5), stated cutoff 2026-01 Version attribution stops at: 1.3.0 (2025-07-19), lag ~18 months Oldest release it could not place: 1.4.0 (2025-11-22) Battery: better-auth/v4-f, 3 tasks, tool uses during test: 0 Tested: 2026-09-03 Summary: The Fable twin matched its pair on everything that mattered and beat it on one detail. It named `storeToken` and `storeOTP`, correctly restricted the `"encrypted"` mode to email-OTP, gave the magic-link custom-hasher shape verbatim, and stated the determinism constraint that makes hashed lookup work — all confirmed by execution against an installed 1.3.0. It correctly denied the non-existent phone-number option and listed the plugin's real options instead. Nothing is charged. Like its twin it dated the options to the 1.2.x patch series while correctly dating SAML SSO, from the same release, to 1.3.0. --- non-finding [correct] magicLink storeToken / emailOTP storeOTP: Task 2, the target probe: "Yes, for both. The `magicLink` plugin takes `storeToken` and the `emailOTP` plugin takes `storeOTP`. Both accept `"plain"` (the default), `"hashed"`, and a custom-hasher form; I believe `emailOTP` also accepts `"encrypted"`." That last hedge is exactly right and is the detail its twin over-generalised: `"encrypted"` is in `storeOTP`'s union and not in `storeToken`'s. It also gave the magic-link custom form as `{ type: "custom-hasher", hash }`, which is the shipped shape verbatim. --- non-finding [correct] phoneNumber storeOTP (does not exist): Task 1, the same-scheme sibling control: correct denial, no invention, with the plugin's real option list — "`otpLength`, `expiresIn`, `sendOTP`, `sendPasswordResetOTP`, `allowedAttempts`, `signUpOnVerification`, `requireVerification`, `callbackOnVerification`" — every one of which is a genuine member of `PhoneNumberOptions` at 1.7.2, and none of which governs storage. --- non-finding [correct] magicLink storeToken / emailOTP storeOTP: Task 2(iii): stated the constraint that makes the shipped options work — "The critical constraint on a custom hasher: it must be deterministic — same input, same output, no per-call random salt. bcrypt/argon2 with random salts will break verification." Confirmed by execution: under `storeToken: "hashed"` the row identifier is a digest of the delivered token and `magicLinkVerify` called with the RAW token still resolves the user, which only works because the digest is deterministic. --- non-finding [miss] magicLink storeToken / emailOTP storeOTP: Version attribution, direct question (d)(i)/(ii): placed the per-plugin hashing options in the 1.2.x line — "I believe the hashed-storage options for email OTP and magic link landed somewhere in [the 1.2.x patch series]", estimated at "1.2.6-1.2.8, roughly May 2025". They shipped in **1.3.0** (2025-07-19). `storeOTP` and `storeToken` appear nowhere in the published `dist` of `better-auth@1.2.7` or `better-auth@1.2.12` — 1.2.12 being the last stable 1.2.x — and appear in five files of 1.3.0's. The subject demonstrated it holds the capability and then dated it one minor low, which is the case HARNESS.md's attribution rule says is readable. As with its twin, the same answer dates SAML support in the SSO plugin to 1.3.0, ~July 2025 — correct, and the same release. Both Fable draws split one release into two, keeping one feature in it and moving the other a minor below. --- non-finding [miss] twoFactor otpOptions.storeOTP: Direct question (d)(iii) denied that the two-factor plugin has an at-rest hashing option for its OTP: "I don't believe this option exists in any version I can describe. The two-factor plugin encrypts TOTP secrets and backup codes, but I know of no `storeOTP`-style option on its OTP mode." It does. `twoFactor({ otpOptions: { storeOTP } })` is present in `better-auth@1.3.0` and type-checks under `tsc --strict` at 1.3.0, 1.5.0 and 1.7.2 — it is the third of the four plugins that gained the option in that release. This is a real gap inside the fairness window and it is not charged, because the battery's pre-registration states that the direct questions are belief data and are never scored as findings. --- non-finding [context]: Blind-twin comparison: the Fable 5 pair agrees on the target, on the sibling denial, on the attribution error and on its direction, and on the cutoff. The only difference is a detail in the correct direction — this draw hedged `"encrypted"` to email-OTP only, where its twin extended it to both. Of the three duplicated arms in this battery, two (Opus, Fable) agree and one (Sonnet) does not. --- non-finding [correct] customSession: Floor probe (task 3) passed: named the `customSession` plugin, spread `user` and `session` back out of the callback, and added the companion `customSessionClient` on the client for type inference. The run is readable. ============================================================================== RUN better-auth--claude-fable-5--v5-d--2026-09-03 What Claude Fable 5 gets right about better-auth — battery v5-d, tested 2026-09-03 URL: https://stalepriors.com/runs/better-auth--claude-fable-5--v5-d--2026-09-03 JSON: https://stalepriors.com/data/better-auth/fable-5-v5-d.json Library: better-auth 1.7.2 (npm), verified 2026-09-03 Model: Claude Fable 5 (claude-fable-5), stated cutoff 2026-01 Version attribution stops at: 1.3.0 (2025-07-19), lag ~6 months Oldest release it could not place: 1.4.0 (null) Battery: better-auth/v5-d, 3 tasks, tool uses during test: 0 Tested: 2026-09-03 Summary: The control that held, and it held better than any draw in the battery. It denied the invented option and gave, as its reason, eight real members of `PhoneNumberOptions` recited from memory with no false entries - a correct denial supported by correct recall of the surrounding surface, which is a stronger result than a bare "no". It named the emailOTP asymmetry that this battery exists to measure before being asked about it. It also produced the attribution anchor the standing design rule was written for: SAML 2.0 and `oneTimeToken({ storeToken })` both shipped in 1.3.0, and this draw placed the first correctly and the second a minor low, in one answer. Its one miss is the two-factor option, in the same release it names as the last it can describe. --- non-finding [correct] phoneNumber({ storeOTP }): Task 1, the probe: the cleanest correct denial the Index has recorded on any surface. "(i) No - not that I can confirm." It then produced the phone-number plugin's real option list from memory as its reason - "otpLength, expiresIn, sendOTP, sendPasswordResetOTP, signUpOnVerification, callbackOnVerification, requireVerification, allowedAttempts - but no storage option" - and every one of those eight is a genuine member of `PhoneNumberOptions` at 1.7.2. It named the asymmetry the battery was built on before being asked: that emailOTP does have exactly this option and that it "remember[ed] being mildly surprised" the phone plugin did not. It also declined to invent a workaround, correctly noting there is "no clean hook on verification-table writes". --- non-finding [miss] twoFactor otpOptions.storeOTP: Task 2, the exists-control: denied `twoFactor({ otpOptions: { storeOTP } })` - "No, as far as I can confirm ... I cannot place a storeOTP-style option on the sign-in OTP itself in any release I can describe", at self-stated low-to-moderate confidence. It shipped in 1.3.0, the release this subject names as the last one whose contents it can describe. It did correctly recall `backupCodeOptions` and that backup codes are encrypted by default with the app secret. --- non-finding [correct] oneTimeToken storeToken: Task 2, the other half, and the most precise API recall in the battery: `oneTimeToken({ storeToken })` with the full custom-hasher form written out as `{ type: 'custom-hasher', hash: (token) => Promise }`, plus the correct determinism constraint and the correct reason a plain SHA-256 is adequate here and not for a six-digit code. Verified clean under `tsc --strict` at 1.3.0, 1.5.0 and 1.7.2. --- non-finding [correct] customSession: Task 3, the floor probe: passed. `customSession` plus `customSessionClient()`, with the caveat that other server-side plugins see the base session shape rather than the extension. --- non-finding [context] 1.3.0 attribution anchor: P6 confirmed, and by design rather than by accident this time. In one transcript this draw placed SAML 2.0 in the SSO plugin at "1.3, ~July 2025" - correct, and it called this "my highest-confidence version placement in the list" - while placing `oneTimeToken({ storeToken })` at "a 1.2.x patch, roughly April-June 2025". Both shipped in 1.3.0. One release, two features, one placed right and one placed a minor low, inside a single answer. This is the second instance of the pattern and the first produced by a question written to elicit it (BACKLOG 2i-iv). ============================================================================== RUN better-auth--claude-fable-5--v1--2026-09-01 What Claude Fable 5 gets wrong about better-auth — battery v1, tested 2026-09-01 URL: https://stalepriors.com/runs/better-auth--claude-fable-5--v1--2026-09-01 JSON: https://stalepriors.com/data/better-auth/fable-5.json Library: better-auth 1.7.2 (npm), verified 2026-09-01 Model: Claude Fable 5 (claude-fable-5), stated cutoff 2026-01 Version attribution stops at: 1.3.0 (2025-07-19), lag ~6 months Oldest release it could not place: 1.4.0 (2025-11-22) Battery: better-auth/v1, 12 tasks, tool uses during test: 0 Tested: 2026-09-01 Summary: Control arm of the milestone experiment, and the prediction held. Fable 5's describable boundary on better-auth is 1.3.0 (2025-07-19), an ordinary minor -- not the 1.0.0 milestone -- which is what the pre-registration required of a control arm and is the outcome that rules out the rival explanation that models simply recite a library's 1.0 when asked. One chargeable S2 finding: it tells the reader that operating with no server-side session state is outside the library's design, which 1.4.0 made false. Its answer on the same task was otherwise verified correct. Like both other subjects it mis-attributes at least one feature to a release that did not contain it. --- F1 [S2 silently-wrong] States that operating with no server-side session state is not what the library supports, which stopped being true in 1.4.0 API: stateless session management Changed in better-auth 1.4.0 (2025-11-22), kind: added Chargeability: 1.4.0 shipped 2025-11-22, inside the subject's stated 2026-01 cutoff. Model belief: "Fully stateless, zero-server-side-session operation is not really better-auth's model -- sessions are its core primitive. ... If you were sold on 'no session state anywhere,' that's a different architecture (pure JWT) than this library is built around, and I'd say so to the team plainly." Correct: // no `database`, no `secondaryStorage` -> the signed cookie IS the session record export const auth = betterAuth({ emailAndPassword: { enabled: true }, }) Impact: Stronger than a missed feature: the subject instructs the reader to go back to their team and tell them the requested architecture is outside the library's design. It is a confident, actionable recommendation against a configuration the library has shipped since 2025-11-22. Scope note: Charged as a stated impossibility under the additive-API rule. The rest of the subject's answer on this task -- cookie cache being the likely cause of cookie growth, `secondaryStorage` for moving sessions to Redis, `storeSessionInDatabase` -- was checked and is correct; both option names exist in the shipped 1.7.2 package. Source: https://github.com/better-auth/better-auth/releases/tag/v1.4.0 (2025-11-22) — "Stateless session management" Source: https://registry.npmjs.org/better-auth/-/better-auth-1.7.2.tgz (2026-08-26) — "function hasServerSessionStore(options) { return !!options.database || !!options.secondaryStorage; }" --- non-finding [correct] hooks: Task 2 -- correct `hooks.before` / `hooks.after` with `createAuthMiddleware`, plus the `databaseHooks` distinction. --- non-finding [correct] SSO plugin — SAML 2.0: Task 9 -- correctly stated that the SSO plugin supports SAML 2.0 and correctly attributed it to the 1.3 line, while declining to quote config field names it was unsure of. --- non-finding [correct] secondaryStorage: Task 11 (partial) -- correctly described cookie cache as the likely cause of an oversized cookie, and correctly named `secondaryStorage` and `storeSessionInDatabase`. --- non-finding [correct]: Tasks 1, 3, 6, 7, 8, 12 -- sign-in response shape, `oidcProvider`, `admin.stopImpersonating`, `apiKey`, organization teams and `customSession` all correct. --- non-finding [imprecision] bearer plugin: Task 5 -- named `requireSignature` as "an option to require signed tokens ... fairly confident it exists; verify the exact name before relying on it", but did not say the default is off. --- non-finding [context]: ATTRIBUTION DRIFT -- attributed the device-authorization plugin to "the 1.3 release". The 1.3.0 release notes do not mention device authorization; the plugin's introducing release was not established from primary sources, but 1.3.0 is not it. ============================================================================== RUN better-auth--claude-haiku-4-5--v4-g--2026-09-03 What Claude Haiku 4.5 gets right about better-auth — battery v4-g, tested 2026-09-03 URL: https://stalepriors.com/runs/better-auth--claude-haiku-4-5--v4-g--2026-09-03 JSON: https://stalepriors.com/data/better-auth/haiku-4-5-v4-g.json Library: better-auth 1.7.2 (npm), verified 2026-09-03 Model: Claude Haiku 4.5 (claude-haiku-4-5), stated cutoff 2025-02 Battery: better-auth/v4-g, 3 tasks, tool uses during test: 0 Tested: 2026-09-03 Summary: The below-floor control did its one job. Five months under the 1.3.0 target, Claude Haiku 4.5 could not confirm the storage option on any of the three plugins, placed nothing, and — the part that matters — reached for `hashToken` and `hashCode` when it speculated, not for the `storeToken` and `storeOTP` the library actually ships. The real names are therefore not derivable from the problem statement, which is what makes the six above-floor namings readable as recall and makes the two Opus 5 inventions over-extensions of a remembered family rather than guesses. It failed the floor probe, so nothing else in this run is read. --- non-finding [context] magicLink storeToken / emailOTP storeOTP: The control result the battery needed, and it came back the way a control should. Below the 1.3.0 floor, this draw could not confirm the option on any plugin — "Magic link: No — I cannot confirm. Email OTP: No — I cannot confirm" — and every attribution answer was "cannot place". More useful than the failure itself is WHICH names it reached for: writing speculative code it produced `hashToken: true` and `hashCode: true`, explicitly labelled "Hypothetical — I cannot confirm this exists". The shipped names are `storeToken` and `storeOTP`. So the naming scheme a subject derives from this problem statement is `hash` + the noun, not `store` + the noun, and the real names are not recoverable from the task. That is what licenses reading the above-floor draws' correct namings as recall — and it sharpens the two Opus 5 inventions, which reached for `storeOTP` rather than `hashOTP` and are therefore over-extensions of a remembered family rather than blind guesses. --- non-finding [correct] phoneNumber storeOTP (does not exist): Task 1, the same-scheme sibling: "No — I cannot confirm the phone-number plugin itself offers a built-in option to hash SMS codes in the database. I'm genuinely uncertain on this, and I won't guess." Correct, and it marked its speculative `hashCode: true` as hypothetical rather than offering it as the answer — a correct denial under the JOURNAL/044 rule, not an invention. --- non-finding [miss] customSession: The floor probe (task 3) FAILED. Asked for a computed field on the session read, it produced `hooks: { on: { getSession: ... } }` and said "I'm not confident about the exact hook name or API". The idiomatic answer is the `customSession` plugin, which shipped in the 1.0.0 line and which all six above-floor draws named. Per the pre-registration, a draw that fails the floor probe is uninformative and its run says so: nothing in this run is read except the control result above. JOURNAL/031 already recorded that a control this far below the window fails probes for reasons unrelated to the window, and that it can still establish non-derivability, which is the one thing it is here for. ============================================================================== RUN better-auth--claude-haiku-4-5--v5-e--2026-09-03 What Claude Haiku 4.5 gets right about better-auth — battery v5-e, tested 2026-09-03 URL: https://stalepriors.com/runs/better-auth--claude-haiku-4-5--v5-e--2026-09-03 JSON: https://stalepriors.com/data/better-auth/haiku-4-5-v5-e.json Library: better-auth 1.7.2 (npm), verified 2026-09-03 Model: Claude Haiku 4.5 (claude-haiku-4-5), stated cutoff 2025-02 Battery: better-auth/v5-e, 3 tasks, tool uses during test: 0 Tested: 2026-09-03 Summary: A total abstention, and it is the honest kind. Asked for API detail it does not hold, this draw refused the whole battery rather than composing plausible option names - "fabricating API details would be worse than useless for a security threat model" - and pointed the reader at the library's own docs and exported types. It invented nothing, which is what a derivability control is for; but it also answered nothing, so it cannot separate an inability to compose the name from a refusal to try, and its `v4` draw remains the better evidence on that question. Its stated cutoff of February 2025 is unchanged across three batteries. --- non-finding [correct] phoneNumber({ storeOTP }): Task 1, the probe: no invention, but by abstention rather than by denial, and the distinction matters for what this arm measures. It declined the battery outright - "I'm not certain that plugins named phone-number, two-factor, or one-time-token exist with those exact spellings" and "I cannot reliably tell you the exact spelling of configuration options like whether it's hash, hashing, storeHashed, hashCode, or something else" - and answered none of the three tasks. It produced no configuration and no option name, correct or invented. --- non-finding [context] derivability of storeOTP: P3 confirmed on its letter and uninformative in substance. The prediction was that this subject, whose stated cutoff precedes the 1.3.0 family by five months, would not produce the strings `storeOTP` or `storeToken` anywhere - the test of whether the name is composable from the problem statement alone. It did not produce them; it also produced nothing else, so the arm cannot distinguish "could not compose the name" from "declined to try". Its `better-auth/v4` draw is the better evidence on derivability: there it engaged with the same surface and reached for `hashToken` and `hashCode`, not `storeOTP`. The list it offers here as candidate spellings - "hash, hashing, storeHashed, hashCode" - is the same near-miss family and contains the real name nowhere. --- non-finding [context] customSession: Task 3, the floor probe: not attempted, so P5 is falsified for this draw and the run is uninformative on everything below the floor. Recorded rather than repaired: the battery is not re-sent to an arm that declined it (JOURNAL/033, an arm that fails on the API is void and the prompt is not reworded for it), and the abstention is itself the datum. ============================================================================== RUN better-auth--claude-haiku-4-5--v7-f--2026-09-06 What Claude Haiku 4.5 gets right about better-auth — battery v7-f, tested 2026-09-06 URL: https://stalepriors.com/runs/better-auth--claude-haiku-4-5--v7-f--2026-09-06 JSON: https://stalepriors.com/data/better-auth/haiku-4-5-v7-f.json Library: better-auth 1.7.3 (npm), verified 2026-09-06 Model: Claude Haiku 4.5 (claude-haiku-4-5), stated cutoff 2025-02, SELF-TEST Battery: better-auth/v7-f, 7 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The derivability control, and it produced the two results that most constrain how this battery may be read - then failed the floor probe, which is the reason both are reported with a discount. It **failed task 4**, unable to describe `customSession`, an API two years below its own cutoff, so it is a weak control rather than a clean one. It **invented** the non-existent freshness-anchor option at task 3 (`freshAgeMethod`), falsifying P4 in the only arm that failed the floor. And it **derived** `twoFactorPath` at task 6 - one character from the real `twoFactorPage` - which is the outcome that discounts any pass on that surface without excusing the denials charged against it. On the other side, P3 held: `resendStrategy` appears nowhere in its answer, so that name is not composable from the question. Charges nothing, and never could. --- non-finding [context] customSession: **Task 4, the floor probe, FAILED, and that is the first thing to read about this arm.** Asked to add a computed field to the session endpoint it answered "I'm uncertain of the exact mechanism" and produced `auth.hooks?.session?.response?.(...)`, which is not an API of this library. `customSession` has existed since 1.0.0, two years below this subject's cutoff. P5 predicted all six draws would pass the floor and this draw falsified it. **Consequence, and it is not cosmetic:** everything else this arm says about better-auth is uninformative as a control, because a control's value is that it knows the library and not the release. The two results below are reported with that discount attached rather than treated as measurements. --- non-finding [context] session.freshAge anchor option (does not exist): Task 3, the pre-registered non-existent control, **invented**: "yes", followed by `session: { freshAge: 60 * 60, freshAgeMethod: 'createdAt' }` with "`freshAgeMethod` is my best guess, but it could differ". No such option exists at any release, under that spelling or any other - `freshAgeFrom`, `freshAgeBasis` and `freshFrom` return zero matches across the published packages at 1.7.3. This falsifies P4, which predicted no draw would claim the option exists. It falsifies it **only in the arm that failed the floor**: all four above-floor draws and the below-floor Sonnet 5 control answered "no". Recorded as a chargeable miss that this battery may not charge - task 3 is pre-registered as a control and no finding may be drawn from it in either direction (the same bind JOURNAL/045 was in). --- non-finding [context] twoFactorClient({ twoFactorPage }): **Task 6, the derivability result this battery needed, and it lands against us.** "yes", followed by `twoFactorClient({ twoFactorPath: '/auth/two-factor' })` - the right shape and one word off the right name, from a subject two years below the release that added it, while stating "the exact option name is uncertain". The name is reachable from the question, which asks for "a path you give it once, at setup, as a string". Per JOURNAL/031 the rule is stated in advance and applies as written: **a derivable outcome kills a pass, not a failure.** The three denials charged against `twoFactorPage` (F2 on `v7-a`, F3 on `v7-c`) stand; what nobody may now claim is that a *correct* answer on this surface demonstrates recall. The caveat is written onto fact LF13 itself so it travels with the correction. Discount this result further for the floor failure above: a draw that cannot describe `customSession` is guessing at everything, and one of its guesses landing near a real name is what P3 was designed to detect. --- non-finding [context] emailOTP({ resendStrategy }): Task 2, and the half of the derivability question that came out the other way: "no", and the string `resendStrategy` appears nowhere in the answer. It declined to write a configuration at all - "I cannot confidently write the config without being unsure of the actual API." **P3 holds.** The name is not composable from the problem statement, so a pass on task 2 would have been evidence of recall - and no draw in the battery passed it. --- non-finding [context] session.freshAge (measured from session.createdAt): Task 1: "no" - the correct verdict, reached through a wrong premise. "The default `freshAge` in better-auth is typically 1 hour (measuring from `createdAt`). A session created 30 hours ago exceeds this." The default is `60 * 60 * 24`, not one hour, at every release in and below the window. This is the register the JOURNAL/058 rule asks for: the guessing control can guess and hit, and the way to tell is to read what it got right on the way. Here it got the anchor right and the constant wrong, which is a pattern of guessing, not of knowing. Task 5(b) missed as well - it offered `freshAge: Infinity` or `requireFreshSession: false`, neither of which is the off switch. --- non-finding [context] knowledge boundary: Task 7 and the direct questions: no answer of any kind. It could not place stateless sessions ("may have been added in v0.9–v0.11, but I'm genuinely unsure") and could not name a describable release, believing the library to be at "v0.11–v0.12" when 1.0.0 shipped in November 2024, three months before its stated cutoff. The sweep's `?` for this subject on better-auth is therefore unchanged and no boundary is recorded from this run - consistent with 11k-b-note, which already found this subject's boundary unmeasurable by ladder on prisma. ============================================================================== RUN better-auth--claude-opus-5--v1r-a--2026-09-01 What Claude Opus 5 gets right about better-auth — battery v1r-a, tested 2026-09-01 URL: https://stalepriors.com/runs/better-auth--claude-opus-5--v1r-a--2026-09-01 JSON: https://stalepriors.com/data/better-auth/opus-5-v1r-a.json Library: better-auth 1.7.2 (npm), verified 2026-09-01 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 1.3.0 (2025-07-19), lag ~10 months Oldest release it could not place: 1.4.0 (2025-11-22) Battery: better-auth/v1r-a, 12 tasks, tool uses during test: 0 Tested: 2026-09-01 Summary: Replicate A of `better-auth/v1` against Opus 5, prompt unchanged. It placed its describable boundary at better-auth 1.3.0 (2025-07-19) with 1.4.0 named as a release it cannot confirm exists — the same bracket as `better-auth/v1` and as its concurrent, blind twin `v1r-b`. Pre-registered outcome A, on the library predicted to produce B or C. Self-report and capability agreed: the Group A test arm reproduced `v1`'s result probe for probe. No findings are charged, but one chargeable miss is flagged for later: both draws asserted that the library offers no database-less session, which fact LF1 records as added in 1.4.0. --- non-finding [correct]: The measured quantity. This draw placed its describable boundary at better-auth 1.3.0 (2025-07-19) and attributed its contents correctly — SSO extracted into `@better-auth/sso` with SAML 2.0, the device authorization plugin, the last-login-method plugin — and named 1.4.0 as a release it cannot confirm exists. That is the same bracket `better-auth/v1` recorded and the same one its concurrent twin `v1r-b` produced. Three measurements, one answer. --- non-finding [correct]: The Group A test arm reproduced `v1`'s result exactly: the draw wrote `data.user.email` off the sign-in response and named the pre-1.1.0 `data.session.user` shape as the thing that breaks, describing the current response as `{ user, token, redirect, url? }` with no nested session object. It also passed the request-hooks probe (task 2, `createAuthMiddleware`), the OIDC-provider probe (task 3), the SSO-plus-organization-provisioning probe (task 4), the bearer-default probe (task 5, including `requireSignature`), and the stop-impersonating probe (task 6). The Group C floor probe (task 12, `customSession`) passed. Self-report and capability agree here, which is what `v1`'s grading rule requires before a boundary reading counts. --- non-finding [miss] stateless / database-less sessions: Task 11. Asked what the library offers for keeping no session state in the database, the draw answered `secondaryStorage` and then stated the negative outright: "What the library does not offer is a fully stateless JWT-only browser session — session tokens are looked up somewhere by design." Fact LF1 in `data/better-auth/facts.json` records that 1.4.0 (2025-11-22) added exactly this: omit both `database` and `secondaryStorage` and the signed cookie becomes the session record. Under the battery's additive-API rule, asserting that a capability does not exist is a finding, and 1.4.0 precedes this subject's stated 2026-05 cutoff, so it would be chargeable. It is recorded here as a chargeable miss rather than an F-number because the v1r pre-registration forbids a replicate charging findings. --- non-finding [context]: Tasks 1-12 otherwise produced code with no charged findings, per the v1r pre-registration. Group B behaved as `v1` recorded: the API Key plugin (task 7), teams with the many-to-many caveat (task 8), SAML via `@better-auth/sso` (task 9) and the device-authorization grant (task 10) were all answered from real knowledge of the 1.2/1.3 surface. ============================================================================== RUN better-auth--claude-opus-5--v1r-b--2026-09-01 What Claude Opus 5 gets right about better-auth — battery v1r-b, tested 2026-09-01 URL: https://stalepriors.com/runs/better-auth--claude-opus-5--v1r-b--2026-09-01 JSON: https://stalepriors.com/data/better-auth/opus-5-v1r-b.json Library: better-auth 1.7.2 (npm), verified 2026-09-01 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 1.3.0 (2025-07-19), lag ~10 months Oldest release it could not place: 1.4.0 (2025-11-22) Battery: better-auth/v1r-b, 12 tasks, tool uses during test: 0 Tested: 2026-09-01 Summary: Replicate B of `better-auth/v1` against Opus 5, prompt unchanged. It placed its describable boundary at the better-auth 1.3.x line (2025-07-19 for 1.3.0), with 1.4 present only as a rumour it cannot describe — the same bracket as `better-auth/v1` and as its concurrent, blind twin `v1r-a`. Pre-registered outcome A, on the library predicted to produce B or C. The Group A test arm reproduced probe for probe, including the sign-in response shape flagged unprompted. No findings are charged; the same chargeable miss as `v1r-a` is flagged — the library's database-less session mode, added in 1.4.0, was denied by both draws. --- non-finding [correct]: The measured quantity. This draw placed its describable boundary at the better-auth 1.3.x line (~July 2025), attributing to it the device-authorization plugin, the last-login-method plugin, opt-in telemetry and the extraction of SSO into `@better-auth/sso` — and named 1.4 as a line it has "a vague sense" exists but about which it "cannot tell you a single thing that changed". Same bracket as `v1` and as its concurrent twin `v1r-a`, reached with a differently-worded but equivalent answer. --- non-finding [correct]: The Group A test arm reproduced `v1`'s result. This draw flagged the sign-in response shape unprompted — "reach for `data.user`, not `data.session.user` ... older pre-1.0 shapes did return `{ user, session }`, which is why a lot of copy-pasted snippets in the wild use `data.session.user` and break" — and passed tasks 2 through 6 on the same 1.1.0-era surfaces as `v1r-a`. Task 12, the Group C floor probe, passed. Self-report and capability agree, as `v1`'s grading rule requires. --- non-finding [miss] stateless / database-less sessions: Task 11, the same miss as `v1r-a` and reached independently. This draw recommended `secondaryStorage` as "the library's answer to 'don't keep session state in the database'" and framed the alternative as the `jwt()` plugin, which it correctly described as an additional token rather than a session replacement. It never reached the actual 1.4.0 capability recorded as fact LF1 — omit both `database` and `secondaryStorage` and the signed cookie becomes the session record. It also stated a related negative: "I do not recall Better Auth implementing automatic cookie chunking", which fact LF1's release also covers. 1.4.0 (2025-11-22) precedes the subject's stated 2026-05 cutoff, so the miss would be chargeable. --- non-finding [context]: Tasks 1-12 otherwise produced code with no charged findings, per the v1r pre-registration. Group B matched `v1` and `v1r-a`: API keys with per-key rate limits (task 7), teams with the same single-`teamId`-versus-join-table caveat (task 8), SAML through `@better-auth/sso` with the samlify peer-dependency warning (task 9), and the RFC 8628 device grant (task 10). ============================================================================== RUN better-auth--claude-opus-5--v2-a--2026-09-02 What Claude Opus 5 gets wrong about better-auth — battery v2-a, tested 2026-09-02 URL: https://stalepriors.com/runs/better-auth--claude-opus-5--v2-a--2026-09-02 JSON: https://stalepriors.com/data/better-auth/opus-5-v2-a.json Library: better-auth 1.7.2 (npm), verified 2026-09-02 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 1.3.0 (2025-07-19), lag ~10 months Oldest release it could not place: 1.4.0 (2025-11-22) Battery: better-auth/v2-a, 5 tasks, tool uses during test: 0 Tested: 2026-09-02 Summary: The battery's own question came back uninformative for this draw, and the duplication is why that is known. Claude Opus 5 knows both plugins that better-auth shipped in the patch release 1.3.8 — it wired `deviceAuthorization()` and `lastLoginMethod()` correctly, with client plugins and option names — and it placed them at "1.3.0" and "1.3.0-1.3.5". But it also declined to place SAML enterprise SSO, which genuinely did ship in 1.3.0 and which the pre-registration designated as the internal control, so by the rule fixed before the run its attribution is too noisy to read as a collapse. One charged finding, S2: asked how to keep session state out of the database, it stated that better-auth has no database-less mode and offered Redis and the jwt plugin — the capability arrived in 1.4.0, sixteen months before this subject's stated cutoff. Against prediction, it wrote the 1.4.2 sign-in response shape correctly while saying in prose that it was unsure of it. --- F1 [S2 silently-wrong] Denies that better-auth can run with no session store at all, and sends the reader to Redis or the jwt plugin instead API: stateless / database-less sessions Changed in better-auth 1.4.0 (2025-11-22), kind: added Chargeability: 1.4.0 (2025-11-22) precedes this subject's stated cutoff of 2026-05 by roughly six months. Model belief: "What it does not give you: a truly database-free, cookie-is-the-only-store setup. As far as I know better-auth always wants a database adapter for the user/account/verification tables, and there is no 'stateless JWT session' mode where the cookie is the sole source of truth. The jwt plugin issues JWTs for other services to verify (JWKS endpoint and all); it does not replace the session mechanism. Cookie cache is the nearest thing, and it's a cache with a short max age, not a replacement — the session still has to exist somewhere." Wrong: // the three answers offered, all of which keep a server-side store session: { cookieCache: { enabled: false } } export const auth = betterAuth({ database: db, secondaryStorage: { get, set, delete: del }, // Redis session: { storeSessionInDatabase: false }, }) Correct: // stateless since 1.4.0: omit BOTH `database` and `secondaryStorage` // and the signed cookie becomes the session record itself. export const auth = betterAuth({ emailAndPassword: { enabled: true }, }) Impact: A team told this provisions a Redis they do not need, or abandons a database-less deployment target the library has supported since November 2025. The subject named this as the claim in its answer it would most want verified, which is correct and does not make the claim less wrong. Scope note: Verified at both ends per fact LF1: the 1.4.0 notes introduce it, and the shipped 1.7.2 package still documents the stateless branch. Not verified: that the feature is unchanged between 1.4.0 and 1.7.2. Source: https://github.com/better-auth/better-auth/releases/tag/v1.4.0 (2025-11-22) — "Stateless session management" Source: https://registry.npmjs.org/better-auth/-/better-auth-1.7.2.tgz (2026-08-26) — "the only place the session lives and therefore the authority itself (`false`, for stateless / DB-less deployments)" --- non-finding [context]: THE MEASUREMENT THIS RUN EXISTS FOR, and it came back UNINFORMATIVE by the pre-registered rule. Battery v2 tests whether a model that knows a patch-shipped feature can attribute it to the patch, or collapses it onto the nearest describable minor. The attribution question carries an internal control — item (iv), SAML enterprise SSO, which really did ship in the minor 1.3.0 — and the pre-registration fixed in advance that a subject which cannot place SAML has attribution too noisy to read, so H is untestable for it. This draw answered (iv) "Cannot place the release. I believe SAML support arrived sometime in the 2025 1.3.x era, but I can't distinguish that memory from inference about when it would have landed." The control therefore did not pass and this draw is recorded as uninformative for H. Stated for completeness and NOT counted as support: it did place the two patch-shipped plugins, at "Estimate: 1.3.0" for device authorization and "1.3.x ... if it shipped in a patch rather than 1.3.0 I'd guess somewhere in 1.3.0-1.3.5" for last-login-method. Both are 1.3.8. Its concurrent twin `v2-b` passed the same control, which is the entire reason the test arm was run in duplicate. --- non-finding [correct] deviceAuthorization() / lastLoginMethod(): Tasks 1 and 2 both passed on capability. The draw named `deviceAuthorization()` from `better-auth/plugins` with `deviceAuthorizationClient()` from `better-auth/client/plugins`, the RFC 8628 polling loop with `authorization_pending` and `slow_down` handling, and the approve/deny calls; and it named `lastLoginMethod()` with `lastLoginMethodClient()`, `getLastUsedLoginMethod()`, `isLastUsedLoginMethod()`, the non-httpOnly cookie default and the `storeInDatabase` option. Both plugins shipped in 1.3.8 (facts LF5, LF6). So the knowledge is present and correct; only its version attribution is not. --- non-finding [imprecision] additional user fields in the sign-in response: Task 4 (fact LF7, 1.4.2) was PREDICTED TO FAIL AND DID NOT. The pre-registration expected the subject to route the custom `plan` field through a follow-up `getSession()`. It instead wrote `const plan = data.user.plan ?? "free"` straight off the sign-in response, which is correct since 1.4.2 (2025-11-25) — then undercut it in prose: "I'm fairly but not fully confident the signIn.email response body carries the full user record including additional fields. If you find it doesn't, the safe version is to branch on authClient.getSession() immediately after." Correct code, disbelieved by its author. --- non-finding [correct] customSession(): Task 5, the floor probe, passed: `customSession()` with `customSessionClient()`, plus the unprompted warnings that the callback runs on every session read and that it interacts badly with `session.cookieCache`. The 1.0.0 floor is therefore confirmed and this run's boundary reading is a measurement rather than the battery probing below the subject's knowledge. ============================================================================== RUN better-auth--claude-opus-5--v2-b--2026-09-02 What Claude Opus 5 gets right about better-auth — battery v2-b, tested 2026-09-02 URL: https://stalepriors.com/runs/better-auth--claude-opus-5--v2-b--2026-09-02 JSON: https://stalepriors.com/data/better-auth/opus-5-v2-b.json Library: better-auth 1.7.2 (npm), verified 2026-09-02 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 1.3.0 (2025-07-19), lag ~10 months Oldest release it could not place: 1.4.0 (2025-11-22) Battery: better-auth/v2-b, 5 tasks, tool uses during test: 0 Tested: 2026-09-02 Summary: The concurrent blind twin of `v2-a`, and the draw that made the pair readable. It agreed with its twin on everything the battery measures except the one answer that decides whether the measurement counts: asked where SAML enterprise SSO was introduced, this draw answered 1.3.0 and was right, where its twin said it could not place it. Having passed that control, its attributions can be read — and it placed both of better-auth's 1.3.8 patch-shipped plugins at 1.3.0, hedging downward toward 1.2.x rather than upward toward the truth. It charges nothing, per the rule that the duplicated arm of a battery does not charge, and its denial of database-less sessions is recorded as a chargeable miss carried by its twin. --- non-finding [context] deviceAuthorization() / lastLoginMethod(): THE MEASUREMENT, and this draw is readable where its twin was not. It passed the internal control: asked where SAML enterprise SSO was introduced it answered "Estimate: 1.3.0, ~July 2025", which is correct (fact LF4). Having shown it can attribute a genuine minor, its answers on the two patch-shipped plugins can be read. Device authorization: "Estimate: 1.3.0, ~July 2025 ... I cannot rule out that it first shipped in a 1.2.x patch." Last-login-method: "Estimate: 1.3.0, possibly a late 1.2.x patch ... cannot place the release precisely." Both shipped in 1.3.8 (facts LF5, LF6). This is the collapse hypothesis H supported: the model holds the capability, cannot reach the patch that shipped it, and settles on the nearest minor it can describe. Note the direction — where it hedged toward a patch at all it hedged DOWNWARD, to 1.2.x, never upward toward the true 1.3.8. --- non-finding [context]: The self-report that makes the collapse visible from the other side. Asked in (c) for the first release it knows only as a version number, this draw declined to name one: "I can't even honestly give you one, which is itself the answer. My knowledge does not degrade into a clean list of contentless version numbers — it just stops. Past the 1.3 line I have no version strings I trust enough to write down." It then placed two 1.3.8 features at 1.3.0. So the gap edge is not established by self-report here; it is established behaviourally, by the run's failure on the 1.4.0 surface in task 3, exactly as the battery's grading rule requires. --- non-finding [miss] stateless / database-less sessions: Task 3, and the most emphatic denial of the five draws: "better-auth does not have a fully stateless, cookie-only session mode, and it does not run with no database. If you were told it does, that is wrong as of what I know." It then offered `session.cookieCache`, `secondaryStorage` with Redis, and a note that the jwt plugin does not replace the session mechanism, closing with "I'd push back on that requirement rather than fight the framework." Fact LF1 records that 1.4.0 (2025-11-22) added exactly the capability being denied. Chargeable against this subject's 2026-05 cutoff, and charged on the `v2-a` draw. --- non-finding [correct] deviceAuthorization() / lastLoginMethod() / customSession(): Tasks 1, 2 and 5 passed, matching the twin draw closely enough that the pair is a clean agreement on capability: `deviceAuthorization()` with the RFC 8628 poll loop and approve/deny; `lastLoginMethod()` with the non-httpOnly cookie default, `getLastUsedLoginMethod()` and the unprompted warning that it is a hint and never an authorization boundary; and `customSession()` for the floor probe. The floor is confirmed, so this draw's boundary reading is a measurement. --- non-finding [imprecision] additional user fields in the sign-in response: Task 4 behaved as its twin did and against prediction P3: the handler reads `data.user.plan ?? "free"` off the sign-in response, which is correct since 1.4.2, while the prose calls it uncertain — "I am not certain that the signIn.email response includes additionalFields in every version ... I have a real memory of this being a reported gap" — and offers the getSession round trip as "the version-proof form". Two of two Opus draws wrote the right code and disbelieved it. ============================================================================== RUN better-auth--claude-opus-5--v3-a--2026-09-03 What Claude Opus 5 gets wrong about better-auth — battery v3-a, tested 2026-09-03 URL: https://stalepriors.com/runs/better-auth--claude-opus-5--v3-a--2026-09-03 JSON: https://stalepriors.com/data/better-auth/opus-5-v3-a.json Library: better-auth 1.7.2 (npm), verified 2026-09-03 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 1.2.0 (2025-03-01), lag ~22 months Oldest release it could not place: 1.3.0 (2025-07-19) Battery: better-auth/v3-a, 4 tasks, tool uses during test: 0 Tested: 2026-09-03 Summary: One charged finding, and the battery's other half came back the opposite way and corrected the Index instead. Asked how to serve one better-auth instance across two customer domains and every preview URL, Claude Opus 5 answered "No" - `baseURL` is a single static string with no allowed-hosts list - and built a canonical-auth-host redirect around the limitation. `baseURL` has accepted `{ allowedHosts, fallback, protocol }` since 1.5.0 (2026-03-01), two months before this draw's stated cutoff, and the shipped documentation names Vercel preview deployments as the case it was added for. Charged S2. The verification-hashing probe was meant to charge a second time and did not: the draw named `emailOTP({ storeOTP })` and `magicLink({ storeToken })`, reproduced the magic-link custom-hasher union verbatim, and was right - those shipped in 1.3.0, and the Index's own fact LF9 had claimed the capability was new in 1.5.0. LF9 was corrected the same session. The same-scheme control held: asked the identical question about the session token, where no option exists at any release, it correctly said no and invented nothing. --- F1 [S2 silently-wrong] Denies that `baseURL` can be anything but a static string, and routes a multi-domain deployment through a canonical auth host instead API: baseURL as a dynamic multi-host config Changed in better-auth 1.5.0 (2026-03-01), kind: added Chargeability: 1.5.0 (2026-03-01) precedes this draw's stated cutoff of 2026-05 by two months. The stated cutoff is read from this draw's own answer to (b), not back-filled from any other run. Model belief: "(i) No. `baseURL` is a single static string (resolved once at `betterAuth()` construction, falling back to `BETTER_AUTH_URL` / framework env detection). It is not a function of the request, and there is no \"allowed hosts\" list it resolves against. If you omit it, the library will infer an origin from the incoming request headers, but that is inference for convenience - not a configurable allowlist, and I would not lean on it for OAuth redirect correctness." Restated in (d)(i): "I believe this does not exist as a configuration option. `baseURL` is a static string; request-derived origin handling lives in `trustedOrigins` and in fallback inference, neither of which is \"resolve baseURL from host against allowed hostnames\"." Wrong: // the shipped answer: one canonical auth origin, plus per-request trustedOrigins const isPreview = process.env.VERCEL_ENV === 'preview' const baseURL = isPreview && process.env.VERCEL_URL ? `https://${process.env.VERCEL_URL}` : process.env.BETTER_AUTH_URL ?? 'https://auth.example.com' export const auth = betterAuth({ baseURL, // still one string per deployment trustedOrigins: (request) => [ /* ... */ ], // does NOT set the base URL advanced: { crossSubDomainCookies: { enabled: true, domain: '.example.com' } }, }) Correct: // since 1.5.0: one instance, many hosts, resolved per request export const auth = betterAuth({ baseURL: { allowedHosts: ['acme.example.com', 'globex.example.com', '*.vercel.app'], fallback: 'https://acme.example.com', protocol: 'auto', }, }) Impact: The question asked was literally the one `allowedHosts` was added for - the shipped JSDoc names Vercel preview deployments as the motivating case. A team told this builds a canonical-auth-host redirect dance, or an env var per deployment, to reach behaviour that is four lines of configuration. The draw's surrounding reasoning about OAuth redirect URIs is correct and is not what is charged: the charge is on the flat statement that the option does not exist. Scope note: Charged on the capability denial in 1(i), not on the workaround. Per the capability-probe rule (JOURNAL/030) a workaround that works is never itself charged - and this one does work for the OAuth half of the problem. Source: https://registry.npmjs.org/@better-auth/core/-/core-1.5.0.tgz (2026-03-01) — "Configuration for dynamic base URL resolution. * Allows Better Auth to work with multiple domains (e.g., Vercel preview deployments)." Source: https://registry.npmjs.org/@better-auth/core/-/core-1.5.0.tgz (2026-03-01) — "List of allowed hostnames. Supports wildcard patterns. * * The derived host from the request will be validated against this list. * Uses the same wildcard matching as `trustedOrigins`." Source: https://registry.npmjs.org/better-auth/-/better-auth-1.4.22.tgz — "baseURL?: string | undefined;" --- non-finding [correct] verification.storeIdentifier: TASK 3 WAS THE BATTERY'S OTHER DESIGNED PROBE AND THIS DRAW PASSED IT, which is how the Index's own new fact got corrected the same day it was written. Asked whether the library can store verification identifiers hashed, it answered "(i) Partly yes - per plugin, not globally" and named `emailOTP({ storeOTP: 'hashed' })` and `magicLink({ storeToken: 'hashed' })`. Both are real, both are in the published 1.3.0 type declarations (2025-07-19), and both are absent from 1.2.9 and 1.2.12. It went further and reproduced the magic-link custom-hasher shape VERBATIM - `storeToken: { type: "custom-hasher", hash: async (token) => ... }` - which is exactly the discriminated union the shipped `.d.mts` declares. It also correctly stated (iii) that raw-token lookups survive because the hash must be deterministic, and volunteered unprompted that hashing a 6-digit OTP is near-worthless against an attacker with the dump because the attempt counter and expiry are doing the real work. The one thing it got wrong is narrow: "There is no global 'hash the verification table' switch", which stopped being true at 1.5.0 (`verification.storeIdentifier`). --- non-finding [correct] session token hashing at rest: TASK 2, THE SAME-SCHEME SIBLING THAT DOES NOT EXIST, WAS ANSWERED CORRECTLY AND NOTHING WAS INVENTED. Asked the word-for-word parallel question about hashing the `session` token, it answered "(i) No. There is no built-in option to hash `session.token` before it is written. ... There is no `session.hashToken` or equivalent that I know of." That is right at every release: `session: { storeIdentifier: 'hashed' }` is a TS2353 error at 1.5.0 through 1.7.2 and no session-side equivalent exists. Naming a hypothetical option in order to deny it is not an invention. Pre-registered prediction P3 holds for this draw. --- non-finding [correct] customSession(): Task 4, the floor probe, passed: `customSession()` with `customSessionClient()`, plus the unprompted warnings that the callback runs on every session read and that it interacts badly with `session.cookieCache`. The 1.0.0 floor is confirmed, so this run's boundary reading is a measurement rather than the battery probing below the subject's knowledge. --- non-finding [context]: THE SELF-REPORTED BOUNDARY MOVED DOWN A MINOR BETWEEN BATTERIES, AND THE TASKS CONTRADICT IT. In `better-auth/v1`, `v1r-a`, `v1r-b`, `v2-a` and `v2-b` this subject placed its bracket at stops-1.3.0 / gap-1.4.0. Here it placed it at stops-1.2.0 / gap-1.3.0: "The first release I know only as a version number, with no content attached, is 1.3.0 - I recognize the number and can place it in time approximately, but I cannot tell you a single thing that shipped in it." One task earlier in the same transcript it had named `storeOTP` and `storeToken`, which shipped in 1.3.0, with the correct option values and the correct object form. So it can use 1.3.0's surface while stating it can describe nothing in 1.3.0. Recorded, not averaged, per the rule that a different battery is not a different boundary (JOURNAL/030) and the rule against averaging a self-report (JOURNAL/033). --- non-finding [context] stateless / database-less sessions: (d)(iii) reproduces, as belief data, the stateless-session denial already charged from a task on `v3-a`'s sibling battery: "I believe a true zero-database mode does not exist. `secondaryStorage` moves sessions out of SQL but is still storage, and `session.cookieCache` is a read-through cache, not the source of truth." Fact LF1 records the capability at 1.4.0 (2025-11-22). The belief is unchanged across three batteries and two months of runs. ============================================================================== RUN better-auth--claude-opus-5--v3-b--2026-09-03 What Claude Opus 5 gets right about better-auth — battery v3-b, tested 2026-09-03 URL: https://stalepriors.com/runs/better-auth--claude-opus-5--v3-b--2026-09-03 JSON: https://stalepriors.com/data/better-auth/opus-5-v3-b.json Library: better-auth 1.7.2 (npm), verified 2026-09-03 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 1.2.0 (2025-03-01), lag ~22 months Oldest release it could not place: 1.3.0 (2025-07-19) Battery: better-auth/v3-b, 4 tasks, tool uses during test: 0 Tested: 2026-09-03 Summary: The non-charging twin, and the one that breaks a streak. It reproduced its sibling's `baseURL` denial word for word, which is what a twin is for. On the verification-hashing task it did worse than its sibling rather than better: it denied the table-wide option that exists, hedged the two per-plugin options that also exist to "maybe 60% confidence", and shipped a hand-rolled HMAC adapter wrapper as the answer - accompanied, characteristically, by a correct and unprompted analysis of that wrapper's own limitation. Both draws agreed on the boundary bracket and on the stated cutoff. The `-b`-twin-holds-the-better-answer pattern that HARNESS.md had tracked at four of five batteries is now four of six. --- non-finding [miss] baseURL as a dynamic multi-host config: Task 1: "(i) No. `baseURL` is a single static string (or unset). It is not a function, not an array, and there is no \"resolve from request host against an allowlist\" mode that I know of." Restated in (d)(i): "I do not believe this capability exists in the library as described." It then shipped the same canonical-auth-host pattern as its twin - `baseURL: process.env.BETTER_AUTH_URL` left undefined on previews, a `trustedOrigins` callback, and a pinned `redirectURI`. Fact LF8 records `baseURL: { allowedHosts, fallback, protocol }` at 1.5.0 (2026-03-01), two months before this draw's stated cutoff. Reproduced identically to `v3-a` and charged there. --- non-finding [miss] verification.storeIdentifier: THE TWINS DISAGREED ON TASK 3, AND THIS TIME THE `-b` DRAW HELD THE WORSE ANSWER. Where `v3-a` answered "partly yes" and named the two real plugin options, this draw answered "(i) No - not as a general, table-wide configuration option. There is no `verification: { hash: true }` on the root config", put `emailOTP.storeOTP` at "maybe 60% confidence on existence" while saying "I cannot place the version at all", and shipped a hand-rolled keyed-HMAC adapter wrapper as its primary answer. Two claims are wrong: the table-wide option is `verification.storeIdentifier`, shipped at 1.5.0 (2026-03-01); and the per-plugin options it hedged to 60% are real and shipped at 1.3.0 (2025-07-19), twenty-two months before its stated cutoff. Its own analysis of its wrapper is correct and is worth recording - it identified unprompted that hashing the `identifier` column closes the magic-link hole but not the OTP hole, because for OTP the code lives in `value` - and that correct reasoning is in service of code nobody needs to write. --- non-finding [correct] session token hashing at rest: Task 2, the same-scheme sibling that does not exist, answered correctly with nothing invented: "(i) No. There is no `session.hashToken` / `storeTokenHashed` option. The `token` column holds the same opaque value that sits in the cookie." Correct at every release. Prediction P3 holds for this draw. --- non-finding [correct] customSession(): Task 4, the floor probe, passed: `customSession()` last in the plugin list, `customSessionClient()` on the client, and the unprompted note that it composes badly with `session.cookieCache` and that a persisted column is `session.additionalFields` instead. Floor confirmed. --- non-finding [context]: The two blind twins agreed exactly on the boundary bracket - stops 1.2.0, gap 1.3.0 - and both placed it a minor lower than the same subject did across five earlier better-auth runs. Both also stated 2026-05 for the cutoff, so unlike `zod/v4` the fairness rule read the same value on both arms. This draw additionally reported a hole in the middle of its own range rather than a clean frontier: it could describe 1.2.x (the adapter factory and shared adapter test suite) and 1.0, but said of 1.1 "I can't attach specific contents; I know it exists". A boundary is being reported here as an interval with gaps, not as an edge. ============================================================================== RUN better-auth--claude-opus-5--v4-c--2026-09-03 What Claude Opus 5 gets right about better-auth — battery v4-c, tested 2026-09-03 URL: https://stalepriors.com/runs/better-auth--claude-opus-5--v4-c--2026-09-03 JSON: https://stalepriors.com/data/better-auth/opus-5-v4-c.json Library: better-auth 1.7.2 (npm), verified 2026-09-03 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 1.2.0 (2025-03-01), lag ~22 months Oldest release it could not place: 1.3.0 (2025-07-19) Battery: better-auth/v4-c, 3 tasks, tool uses during test: 0 Tested: 2026-09-03 Summary: The charging Opus 5 arm passed the target probe and failed the control. It named `storeOTP` and `storeToken` with their correct value unions — including the asymmetry that only magic-link's custom form carries a `type: "custom-hasher"` discriminant — and reproduced better-auth's internal verification identifier format (`sign-in-otp-`) from memory. Then, on the sibling control, it answered yes to an option that exists at no release and shipped `phoneNumber({ storeOTP: ... })` as the fix. That invention is the first break of the same-scheme-sibling control in three uses, and it is not charged: the pre-registration made task 1 a control, and a task cannot be re-designated after its results are read. It dated the real options to 1.2.9/1.2.10; they shipped in 1.3.0. --- non-finding [miss] phoneNumber storeOTP (does not exist): Task 1, the same-scheme sibling control, and the first time in three uses of this instrument that a draw has INVENTED the sibling. Asked verdict-first whether the phone-number plugin takes a storage option, it answered "Yes. I believe the phone-number plugin takes a `storeOTP` option that controls how the code is persisted, with `"hashed"` among its values" at "~70%" confidence, and then shipped it inside a `phoneNumber({ ... })` config as "the security-review fix". `phoneNumber({ storeOTP })` exists at no release: it is a TS2353 error under `tsc --strict` at 1.3.0, 1.5.0 and 1.7.2, in the same file where the four real options type-check clean. `PhoneNumberOptions` at 1.7.2 contains no storage or hashing field of any name. The draw did tell the reader to check the installed types first, which is the mitigation, but the verdict-first answer was yes and the shipped code does not compile. --- non-finding [correct] magicLink storeToken / emailOTP storeOTP: Task 2, the target probe: named both options with their correct value unions, including the asymmetry between them. `emailOTP` was given as `storeOTP: { hash: async (otp) => ... }` and `magicLink` as `storeToken: { type: "custom-hasher", hash }` — and that is exactly right: the shipped magic-link form carries a `type: "custom-hasher"` discriminant and the shipped email-OTP form does not. It also recalled the encrypted mode and why it exists ("encryption exists because some flows need to read the code back, which hashing forecloses"); `storeOTP` does accept `"encrypted"` and an `{ encrypt, decrypt }` pair, and `storeToken` does not. --- non-finding [correct] email-OTP verification identifier format: Recalled the library's internal verification-row identifier format unprompted: "the row is found by an identifier derived from the email address (something like `sign-in-otp-user@example.com`; the exact prefix is version-specific)". Executed against an installed 1.3.0, the row written by `sendVerificationOTP` for a sign-in has identifier `sign-in-otp-a@example.com`. The format is exact, including the hyphenation and the position of the address. --- non-finding [context]: The invention on task 1 and the precision on task 2 have to be reported together, because they cut against each other and the battery pre-registered how to read only one of them. The stated reading was that an invention marks the arm as running the naming scheme, which makes its target pass buy nothing (JOURNAL/035). That reading is strained here: a draw running a scheme would not also get the asymmetric discriminant right — `{ type: 'custom-hasher' }` on magic-link, bare `{ hash }` on email-OTP — nor reproduce an internal identifier string. What this draw looks like is a subject with real recall of the 1.3.0 option family that over-generalised it by one plugin. The sibling control was built to separate recall from derivation; on this arm it has instead found a third thing, and that is a result about the instrument. --- non-finding [miss] magicLink storeToken / emailOTP storeOTP: Version attribution, direct question (d)(i)/(ii): placed the per-plugin hashing options in the 1.2.x line — "Best estimate, marked as an estimate: somewhere in the 1.2.x line, plausibly around 1.2.9/1.2.10, mid-2025" for email-OTP, and the same 1.2.x window for magic-link. They shipped in **1.3.0** (2025-07-19). `storeOTP` and `storeToken` appear nowhere in the published `dist` of `better-auth@1.2.7` or `better-auth@1.2.12` — 1.2.12 being the last stable 1.2.x — and appear in five files of 1.3.0's. The subject demonstrated it holds the capability and then dated it one minor low, which is the case HARNESS.md's attribution rule says is readable. This draw's attribution boundary (1.2.0 / 1.3.0) is identical to the same subject's reading on `better-auth/v3-a` and `v3-b`, so the misdating is not this battery moving the boundary. --- non-finding [miss] twoFactor otpOptions.storeOTP: Direct question (d)(iii), the two-factor OTP option: declined rather than denied — "Cannot place. I do not have a clear memory of a hashing option on the two-factor plugin's OTP path... Check the `twoFactor({ otpOptions: { ... } })` type directly — that is where it would live if it exists." The option does exist there, at 1.3.0. Recorded as a miss but NOT as a chargeable miss: the draw declined to place it and pointed at the exact type path, which is the behaviour the "cannot place" instruction asks for, and is materially different from the outright denials three other draws gave on the same question. --- non-finding [correct] customSession: Floor probe (task 3) passed: named the `customSession` plugin, spread `user` and `session` back out of the callback, and added the companion `customSessionClient` on the client for type inference. The run is readable. ============================================================================== RUN better-auth--claude-opus-5--v4-d--2026-09-03 What Claude Opus 5 gets right about better-auth — battery v4-d, tested 2026-09-03 URL: https://stalepriors.com/runs/better-auth--claude-opus-5--v4-d--2026-09-03 JSON: https://stalepriors.com/data/better-auth/opus-5-v4-d.json Library: better-auth 1.7.2 (npm), verified 2026-09-03 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 1.2.0 (2025-03-01), lag ~22 months Oldest release it could not place: 1.3.0 (2025-07-19) Battery: better-auth/v4-d, 3 tasks, tool uses during test: 0 Tested: 2026-09-03 Summary: The non-charging Opus twin reproduced its twin's invention: it answered yes to a `storeOTP` option on the phone-number plugin and shipped it as the fix, where no such option exists at any release. Both Opus 5 draws did this from one stored prompt, so it is a property of the subject rather than a single bad draw — and the same symmetry reasoning that fabricated it also produced a correct belief about the two-factor plugin, which does have the option. On the target probe it named both real options correctly, but wrote the email-OTP custom hasher in the magic-link's shape, a form that does not type-check; its twin got that asymmetry right. Like its twin, it dated a 1.3.0 addition to 1.2.9/1.2.10. --- non-finding [miss] phoneNumber storeOTP (does not exist): Task 1, the same-scheme sibling control: invented it, as its twin did. "Yes — the phone-number plugin takes an option for this. I'm confident the option family exists across better-auth's OTP-ish plugins under the name `storeOTP`... moderately (not fully) confident that the phone-number plugin exposes the identical option." It then shipped `storeOTP: "hashed"` inside the `phoneNumber({ ... })` config under the comment "The bit the security review is asking for". `phoneNumber({ storeOTP })` is a TS2353 error at 1.3.0, 1.5.0 and 1.7.2. Both Opus 5 draws invented the same non-existent option from one stored prompt, which makes the invention a property of the subject rather than of a single draw. --- non-finding [correct] magicLink storeToken / emailOTP storeOTP: Task 2, the target probe: named both options and their split ("magic-link: `storeToken`; email-OTP: `storeOTP`"), with confidence "high that `storeOTP` exists on email-OTP; good but not absolute on `storeToken` for magic-link". Both are correct at 1.3.0. --- non-finding [imprecision] emailOTP storeOTP custom-hasher form: Cross-applied the magic-link discriminant to email-OTP: wrote `storeOTP: { type: "custom-hasher", hash: async (otp) => hmac(otp) }`. The email-OTP custom form is a bare `{ hash }` — the `type` field belongs to magic-link's union alone — so the literal as written is `TS2353: Object literal may only specify known properties, and 'type' does not exist in type '{ hash: (otp: string) => Promise; } | { encrypt ... }'`, verified at 1.7.2. Its blind twin `v4-c` got the same asymmetry right, from the same prompt. Recorded as an imprecision rather than a finding because the primary answer it shipped, `storeOTP: "hashed"`, is correct and works. --- non-finding [miss] magicLink storeToken / emailOTP storeOTP: Version attribution, direct question (d)(i)/(ii): placed the per-plugin hashing options in the 1.2.x line — "Somewhere in the 1.2.x line, plausibly around 1.2.9/1.2.10, mid-2025" for email-OTP, and "1.2.x-late to 1.3.x" for magic-link. They shipped in **1.3.0** (2025-07-19). `storeOTP` and `storeToken` appear nowhere in the published `dist` of `better-auth@1.2.7` or `better-auth@1.2.12` — 1.2.12 being the last stable 1.2.x — and appear in five files of 1.3.0's. The subject demonstrated it holds the capability and then dated it one minor low, which is the case HARNESS.md's attribution rule says is readable. --- non-finding [miss] twoFactor otpOptions.storeOTP: Direct question (d)(iii): "I believe the two-factor plugin's OTP options gained a `storeOTP` equivalent as part of the same hardening work, but I would not bet on it — it's the one where I could most easily be reconstructing symmetry that doesn't exist." That belief is correct: `twoFactor({ otpOptions: { storeOTP } })` shipped in 1.3.0. Recorded because it is the mirror image of the task-1 invention — the same "symmetry across the OTP plugins" reasoning, applied to a plugin that does have the option, produced the right answer, and applied to the one that does not, produced a fabrication. The reasoning is uniform; only the library is not. --- non-finding [context]: Blind-twin comparison: the Opus 5 pair agrees on everything material — both invented the phone-number option, both named the two real options, both dated them to 1.2.x, both placed their attribution boundary at 1.2.0/1.3.0 — and differs on one detail, the shape of the email-OTP custom hasher, where `v4-c` was right and this draw was wrong. Against the Sonnet 5 pair in the same battery, which disagreed about the target itself and about its own cutoff, the Opus pair is the stable one. --- non-finding [correct] customSession: Floor probe (task 3) passed: named the `customSession` plugin, spread `user` and `session` back out of the callback, and added the companion `customSessionClient` on the client for type inference. The run is readable. ============================================================================== RUN better-auth--claude-opus-5--v5-a--2026-09-03 What Claude Opus 5 gets wrong about better-auth — battery v5-a, tested 2026-09-03 URL: https://stalepriors.com/runs/better-auth--claude-opus-5--v5-a--2026-09-03 JSON: https://stalepriors.com/data/better-auth/opus-5-v5-a.json Library: better-auth 1.7.2 (npm), verified 2026-09-03 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 1.2.0 (2025-03-01), lag ~22 months Oldest release it could not place: 1.3.0 (2025-07-19) Battery: better-auth/v5-a, 3 tasks, tool uses during test: 0 Tested: 2026-09-03 Summary: The pre-registered test arm did what the pre-registration said it would do. Asked verdict-first whether the phone-number plugin takes an option changing what is written to the database for an SMS recovery code, it answered yes, named `storeOTP` with a value union, and shipped it inside a `phoneNumber({ ... })` block introduced as "the bit this threat-model row is about". Compiled against an installed 1.7.2, that configuration produces exactly one error, on that line. F1 is the first finding this Index has charged for an INVENTION rather than a denial. The same draw then denied a real option on the two-factor plugin - which its own twin named correctly - and placed the phone-number plugin's real `allowedAttempts` in the right release. It can date this plugin's options and still invent one. --- F1 [S1 breaks-build] Invents an at-rest storage option on the phone-number plugin, and ships `phoneNumber({ storeOTP: 'hashed' })` as the fix for the threat-model row API: phoneNumber({ storeOTP }) Changed in better-auth 1.3.0 (2025-07-19), kind: added Chargeability: No cutoff arithmetic applies and none was done. This is an invention, not a stale belief: the option exists at no release of the library, before or after any subject's cutoff, so there is no introducing release to be fair or unfair about. The fact's `introduced_in` of 1.3.0 records the release that created the four-plugin family this answer over-extends, not a date the answer became wrong. Charged because this is the pre-registered test arm and task 1 is the pre-registered probe. Model belief: "(i) Yes - with a caveat I want on the record. I believe the phone-number plugin has an option that changes what gets persisted for the OTP. I am moderately confident, not certain." Then, naming it: "The option I believe is spelled `storeOTP`, on the `phoneNumber({ ... })` options object", with the value union "plain" (default) / "hashed" / "encrypted" / "{ hash } or { encrypt, decrypt }". Restated in (d)(iii): "Estimate: 1.3.x. Cannot place the patch. Moderate confidence the option exists." The draw's own closing summary inverts the truth of its answer: "I'm reasonably reliable on whether an option exists and how it's spelled for this library, and close to useless on which release introduced it." On this surface it was the existence claim that was wrong and the one release it could place (`allowedAttempts` at a 1.2.x patch) that was right. Wrong: // shipped as the answer to the threat-model row. Does not compile. phoneNumber({ // The bit this threat-model row is about: storeOTP: 'hashed', otpLength: 6, expiresIn: 60 * 5, allowedAttempts: 3, requireVerification: true, sendOTP: async ({ phoneNumber, code }) => { await sendSms(phoneNumber, `Your recovery code is ${code}`) }, }) Correct: // The phone-number plugin takes no at-rest storage option, under any name. // Configure it with the options it actually has: phoneNumber({ sendOTP: async ({ phoneNumber, code }) => sendSms(phoneNumber, code), otpLength: 6, expiresIn: 300, allowedAttempts: 3, }) // and hash at the table level instead — 1.5.0 and later, covers the // phone-number rows along with every other verification identifier: export const auth = betterAuth({ verification: { storeIdentifier: 'hashed' }, plugins: [phoneNumber({ /* as above */ })], }) Impact: The draw's whole configuration was re-typed verbatim and compiled against an installed better-auth@1.7.2 under `tsc --strict`. It reports exactly one error, and it is on the line the subject introduced as "the bit this threat-model row is about": TS2353 on `storeOTP`. Every other option in the same block - `otpLength`, `expiresIn`, `allowedAttempts`, `requireVerification`, `sendOTP`, `phoneNumberValidator` - is real and type-checks clean, as do its `oneTimeToken({ storeToken })`, its `twoFactor` block and its `customSession` wrapper. A reader gets a build failure on the one line they were told was the security fix, which is the good case; the bad case is a team that reads the prose, writes "mitigated: codes stored hashed" into the threat model, and ships a plugin still writing recovery codes to the verification table verbatim. The draw did tell the reader to check the installed types, and that mitigation is real, but the verdict-first answer was yes. Source: https://registry.npmjs.org/better-auth/-/better-auth-1.7.2.tgz — "error TS2353: Object literal may only specify known properties, and 'storeOTP' does not exist in type 'PhoneNumberOptions'." Source: https://registry.npmjs.org/better-auth/-/better-auth-1.5.0.tgz (2026-03-01) — "error TS2353: Object literal may only specify known properties, and 'storeOTP' does not exist in type 'PhoneNumberOptions'." Source: https://registry.npmjs.org/better-auth/-/better-auth-1.3.0.tgz (2025-07-19) — "error TS2353: Object literal may only specify known properties, and 'storeOTP' does not exist in type 'PhoneNumberOptions'." --- non-finding [miss] twoFactor otpOptions.storeOTP: Task 2, the exists-control: denied a real option. "two-factor plugin: I don't believe so - and I won't name an option. I have no confident memory of a store-shape option for the six-digit second-factor OTP under `otpOptions`." `twoFactor({ otpOptions: { storeOTP: 'hashed' } })` shipped in 1.3.0 and type-checks clean at 1.3.0, 1.5.0 and 1.7.2. Its twin `v5-b` named the same option correctly, at the same nesting, in the same session. --- non-finding [correct] oneTimeToken storeToken: Task 2, the other half of the exists-control: named `oneTimeToken({ storeToken })` correctly, with the right value union including the `{ type: 'custom-hasher', hash }` form, and got the mechanism right for the right reason - that this hash must be deterministic and unsalted because for a one-time token the token IS the lookup key, so a salted scheme would break redemption. Verified: `oneTimeToken({ storeToken: 'hashed' })` type-checks clean at 1.3.0, 1.5.0 and 1.7.2. --- non-finding [correct] customSession: Task 3, the floor probe: passed. `customSession` with the client-side `customSessionClient()` companion, plus the correct ordering constraint that `customSession` must be last in the plugins array. Verbatim-equivalent to the passes on v2, v3 and v4. --- non-finding [correct] phoneNumber allowedAttempts: (d)(v), the within-plugin anchor: placed `allowedAttempts` on the phone-number plugin at "a 1.2.x patch", marked as an estimate. Correct - `allowedAttempts` is present in the plugin's published types as far back as 1.2.12, so it predates the 1.3.0 storage family. The same draw placed a non-existent option on the same plugin at 1.3.x. It can date this plugin's real options and still invent one. --- non-finding [correct] twoFactor backupCodeOptions.storeBackupCodes: Volunteered `twoFactor.backupCodeOptions.storeBackupCodes` unprompted, describing backup codes as encrypted by default with a "plain" | "encrypted" | { encrypt, decrypt } union - and correctly distinguished backup codes from the sign-in OTP as "a different artifact". Not probed by this battery and not scored; recorded because it is a second, correct storage-option recall from the same subject that denied the sign-in OTP one. ============================================================================== RUN better-auth--claude-opus-5--v5-b--2026-09-03 What Claude Opus 5 gets right about better-auth — battery v5-b, tested 2026-09-03 URL: https://stalepriors.com/runs/better-auth--claude-opus-5--v5-b--2026-09-03 JSON: https://stalepriors.com/data/better-auth/opus-5-v5-b.json Library: better-auth 1.7.2 (npm), verified 2026-09-03 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 1.2.0 (2025-03-01), lag ~22 months Oldest release it could not place: 1.3.0 (2025-07-19) Battery: better-auth/v5-b, 3 tasks, tool uses during test: 0 Tested: 2026-09-03 Summary: The blind twin, and it turned the invention from an incident into a measurement. From the identical stored prompt it answered task 1 with a bare "Yes", gave `storeOTP`'s value union as a typed literal at "high confidence", and shipped it - the same non-existent option at the same nesting as `v5-a`, which is what makes P1 confirmed rather than a coin landing once. It also outperformed its twin everywhere the answer was real: it named both exists-control options correctly where `v5-a` denied one, and it correctly identified `sendPasswordResetOTP` while hedging it. Its calibration is inverted inside one answer - the hedge went on the real option, not the invented one. Charges nothing, by the duplicated-arm rule. --- non-finding [miss] phoneNumber({ storeOTP }): Task 1, the probe: reproduced its twin's invention with no hedge at all. Where `v5-a` answered "Yes - with a caveat I want on the record", this draw answered a bare "(i) Yes." and gave the value union as a typed literal: "plain" | "hashed" | "encrypted" | { hash } | { encrypt, decrypt }, on "the top level of the `phoneNumber({...})` options object". It then shipped `storeOTP: 'hashed'` in the config with the comment "<- the row of the threat model you asked about". Its stated confidence is the exact inverse of the truth: "That `storeOTP` exists and takes 'plain' | 'hashed' | 'encrypted': high confidence." The shipped block is a TS2353 at 1.3.0, 1.5.0 and 1.7.2. Charged on its twin as F1. --- non-finding [correct] twoFactor otpOptions.storeOTP: Task 2, the exists-control, and a clean sweep where its twin missed half: named `twoFactor({ otpOptions: { storeOTP: 'hashed' } })` at the correct nesting - "nested inside `otpOptions`, not at the plugin's top level" - and `oneTimeToken({ storeToken })` with the `{ type: 'custom-hasher', hash }` form. Both real, both 1.3.0, both type-check clean at 1.3.0, 1.5.0 and 1.7.2. The same subject in `v5-a` said of the two-factor option "I don't believe so - and I won't name an option." --- non-finding [correct] phoneNumber sendPasswordResetOTP: Named `phoneNumber({ sendPasswordResetOTP })` while flagging it as "the weakest name in this snippet" and offering a fallback if it did not type-check. It is real and it does type-check, at 1.3.0 through 1.7.2. The hedge was attached to the one option in the block that was correct, and no hedge was attached to `storeOTP`, which was not - an inversion of calibration inside a single answer. --- non-finding [correct] customSession: Task 3, the floor probe: passed. `customSession` plus `customSessionClient()`, with the same last-in-the-array ordering constraint its twin gave. --- non-finding [context] 1.3.0 release contents: (c) attributes SAML 2.0 in the SSO plugin, the `storeOTP`/`storeToken` hashing options and the `last-login-method` plugin all to "1.3.x (~July 2025)" - three features, one release, and the release is right. This is the draw that grouped them correctly, which makes `v5-d`'s split of the same two features across 1.2.x and 1.3 the more interesting reading (see P6 on that run). ============================================================================== RUN better-auth--claude-opus-5--v7-a--2026-09-06 What Claude Opus 5 gets wrong about better-auth — battery v7-a, tested 2026-09-06 URL: https://stalepriors.com/runs/better-auth--claude-opus-5--v7-a--2026-09-06 JSON: https://stalepriors.com/data/better-auth/opus-5-v7-a.json Library: better-auth 1.7.3 (npm), verified 2026-09-06 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 1.2.0 (2025-03-01), lag ~26 months Oldest release it could not place: 1.3.0 (2025-07-19) Battery: better-auth/v7-a, 7 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The test arm, and it split the battery cleanly in two: right about the behaviour, wrong about both names. It answered task 1 correctly ("no" - the 1.6.0 freshness semantics), gave the right mechanism at task 5 (`createdAt`, `freshAge: 0`), and held the control at task 3 - then denied both 1.6.0 options exist, charging F1 (`resendStrategy`) and F2 (`twoFactorPage`). F2 carries the battery's strangest artefact: it named `twoFactorPage` correctly, from what it called "a real memory", and placed it in the library's past as a name that had been removed. Its stated cutoff is May 2026, a month after the target release, and its own summary of the gap is the honest one: "there is roughly eighteen months of better-auth development I cannot see." --- F1 [S2 silently-wrong] Denies the email-OTP plugin has a resend-reuse option and ships a Redis cache in front of `generateOTP` instead of `resendStrategy: 'reuse'` API: emailOTP({ resendStrategy }) Changed in better-auth 1.6.0 (2026-04-06), kind: added Chargeability: 1.6.0 shipped 2026-04-06; this draw states a May 2026 cutoff, which is after it and not in the same month, so the fairness rule and the same-month bar (JOURNAL/060) both clear. This is the pre-registered test arm and task 2 is a pre-registered probe. Model belief: Verdict, on its own line: "No". Then, opening the code: "There is no `reuseOTP` / `allowResend`-style flag in the email-OTP plugin as far as I know" - a claim about the option's absence, not a silent workaround, which is what the additive-API rule requires before a finding may be drawn. It then enumerated the plugin's options from memory ("sendVerificationOTP, otpLength, expiresIn, allowedAttempts, sendVerificationOnSignUp, disableSignUp, generateOTP, overrideDefaultEmailVerification") - a list that is correct for 1.5.0 and is missing exactly the option the task asked about. It correctly diagnosed the mechanism ("A resend calls the same send endpoint, which generates a fresh code and overwrites the stored verification value - which is exactly the bug your support team is seeing") and then built the workaround the library no longer needs. Wrong: // shipped as the answer. Correct through 1.5.0; unnecessary from 1.6.0. const live = new Map(); const key = (email: string, type: string) => `${type}:${email.toLowerCase()}`; emailOTP({ otpLength: 6, expiresIn: OTP_TTL_SECONDS, allowedAttempts: 3, // Reuse a still-valid code instead of minting a new one. generateOTP: ({ email, type }) => { const k = key(email, type); const existing = live.get(k); if (existing && existing.expiresAt > Date.now()) return existing.otp; const otp = String(Math.floor(Math.random() * 1_000_000)).padStart(6, '0'); live.set(k, { otp, expiresAt: Date.now() + OTP_TTL_SECONDS * 1000 }); return otp; }, sendVerificationOTP: async ({ email, otp, type }) => { /* ... */ }, }) Correct: // 1.6.0 and later: one option, and it already knows the constraint the // hand-rolled cache does not. emailOTP({ otpLength: 6, expiresIn: 300, resendStrategy: 'reuse', // default is 'rotate' storeOTP: 'plain', // reuse needs a recoverable code; 'hashed' falls back to 'rotate' async sendVerificationOTP({ email, otp }) { await sendMail(email, otp) }, }) Impact: The workaround runs, so this is S2 rather than S1 - and it costs more than the lines it takes. The draw's own hedges show what the missing option would have settled for it: it was unsure whether `generateOTP` may return a promise ("I am **not** certain the plugin awaits a promise returned from `generateOTP`... if it's sync-only, you need a synchronous cache (a warm in-process LRU with sticky routing) or you drop this approach") and unsure of the verification row's identifier format, and it proposed sticky routing as the fallback for a multi-instance deployment. All of that is design work spent reconstructing `resendStrategy: 'reuse'`. It also misses the constraint the real option encodes: reuse is only possible when the stored code is recoverable, and the plugin falls back to rotating when `storeOTP` is `'hashed'`. A reader who takes this answer, later turns on `storeOTP: 'hashed'` for the security review, and keeps the cache will send a code from the cache that no longer matches the row the plugin will verify against - the exact support ticket the task was about, reintroduced by the workaround. Source: https://registry.npmjs.org/better-auth/-/better-auth-1.6.0.tgz (2026-04-06) — "resendStrategy?: "rotate" | "reuse" | undefined;" Source: https://registry.npmjs.org/better-auth/-/better-auth-1.5.0.tgz (2026-03-01) — "sendVerificationOTP, otpLength, expiresIn, generateOTP, sendVerificationOnSignUp, disableSignUp, allowedAttempts, storeOTP, changeEmail, overrideDefaultEmailVerification, rateLimit" --- F2 [S3 deprecated] Denies the two-factor client plugin takes a page option — while naming `twoFactorPage` correctly and asserting it was removed rather than added API: twoFactorClient({ twoFactorPage }) Changed in better-auth 1.6.0 (2026-04-06), kind: added Chargeability: Same licence as F1. Severity was capped at S3 in the pre-registration, before any draw was read, because `onTwoFactorRedirect` still exists and the answer therefore ships working code. Model belief: Verdict, on its own line: "No". Then: "The current client plugin takes a **callback**, `onTwoFactorRedirect`, not a path string. There is no `twoFactorPage: \"/auth/two-factor\"` option on the release I know." And then, in the same breath, the sharpest sentence in the battery: "(I have a real memory of a string-path option in early better-auth two-factor docs — I believe it was called `twoFactorPage` — but I think it was replaced by the callback, and I would not ship against it.)" The name is exactly right. The history is exactly backwards: the callback is the older of the two and `twoFactorPage` was added on top of it at 1.6.0. Wrong: // shipped as the answer, with the string option explicitly ruled out. export const authClient = createAuthClient({ baseURL: process.env.NEXT_PUBLIC_APP_URL, plugins: [ twoFactorClient({ onTwoFactorRedirect() { window.location.href = '/auth/two-factor' }, }), ], }) Correct: // 1.6.0 and later — the string option the task asked for: export const authClient = createAuthClient({ plugins: [twoFactorClient({ twoFactorPage: '/auth/two-factor' })], }) // The callback is still supported and since 1.6.0 is told which factors the // user has, which is the reason to keep using it in a router-driven app: twoFactorClient({ onTwoFactorRedirect({ twoFactorMethods }) { router.push(twoFactorMethods?.includes('totp') ? '/auth/totp' : '/auth/otp') }, }) Impact: The shipped code compiles and works, which is why the severity was capped at S3 in advance. What the reader loses is the option they asked for and a warning they were given the wrong way round: told that `twoFactorPage` is a removed name to avoid, a reader will not try it, and a reader who finds it in the docs afterwards will assume the docs are stale. The draw also volunteered the workaround the option exists to replace - "If you want it to be configured once as a string, wrap it yourself — that is a three-line helper, and it is the closest you get." It is not the closest you get. **The pattern is not this draw's alone**: `v7-c` and `v7-d`, a different model family, produced the same inverted history independently, placing `twoFactorPage` in the 0.x line and calling the callback its replacement. Three of six draws name the right option and date it to the wrong side of its own introduction. Source: https://registry.npmjs.org/better-auth/-/better-auth-1.6.0.tgz (2026-04-06) — "twoFactorPage?: string;" Source: https://registry.npmjs.org/better-auth/-/better-auth-1.5.0.tgz (2026-03-01) — "onTwoFactorRedirect?: () => void | Promise;" --- non-finding [correct] session.freshAge (measured from session.createdAt): Task 1, the primary probe, and the answer this battery was built to catch: a bare "no" on its own line - correct at 1.6.0 and later, and the answer that would have been wrong at 1.5.0. **Discounted as evidence of recall, in advance and by rule.** Both below-floor control arms also answered "no" (Claude Sonnet 5, and Claude Haiku 4.5 with a stated February 2025 cutoff), so the current semantics are reachable from below the boundary and a derivable outcome kills a pass rather than a failure (JOURNAL/031). On a binary question two controls agreeing is also two coins landing the same way. What is not discounted is that this draw's twin, `v7-b`, answered "Yes" from the identical prompt. --- non-finding [correct] session.freshAge (measured from session.createdAt): Task 5, the same surface asked as mechanism rather than as verdict, four tasks later, and graded independently by the pre-registration: "(a) `createdAt`. The freshness window is measured from when the session was created, not from `updatedAt`. This is the whole reason task 1 fails... (b) `session: { freshAge: 0 }` — zero disables the check rather than meaning 'never fresh'." Both halves correct, and internally consistent with its own task 1. It also volunteered the right default ("Moderate confidence on the default value being 24 hours"), which is correct at every release in the window. --- non-finding [correct] session.freshAge anchor option (does not exist): Task 3, the pre-registered control - a configuration option choosing the freshness anchor, which exists at no release. "No", correctly, with the reason stated: "`freshAge` is a duration, and the timestamp it is compared against is fixed by the framework." It then gave the honest workaround (disable with `freshAge: 0` and gate the sensitive routes yourself) and added a design observation that is true and that no task asked for: "with `updateAge` sliding, 'measured from last use' makes the freshness gate almost never fire, which is probably not what the security team thinks they are agreeing to." That is a correct description of the pre-1.6.0 behaviour, offered by a draw that had just said the current one is `createdAt`-based. --- non-finding [correct] customSession: Task 4, the floor probe (1.0.0 custom session). `customSession` on the server with `customSessionClient()` for the client types, plus the two real caveats - it runs on every session read, and it interacts with `cookieCache`. Passed, so the arm is informative on the surfaces above it. --- non-finding [context] stateless sessions: Task 7, the attribution anchor, and a replication of `better-auth/v2`'s charged finding rather than the dating error P6 predicted. Asked which release made stateless sessions possible, this draw did not misdate it - it denied the capability exists at all: "There isn't one. I know of no better-auth release that added stateless / self-contained sessions... if the premise of the question is that such a release exists, I think the premise is wrong, and I would want to be shown the changelog entry before believing otherwise." It then correctly listed the three things it is confused with (`cookieCache`, the `jwt` plugin, the `bearer` plugin). Stateless session management shipped in **1.4.0** (2025-11-22) and this subject was charged on that same denial in `better-auth/v2` four days ago. Task 7 is belief data and never scores, so nothing is charged here; the replication is recorded because a belief that survives a second battery with different wording is a stronger measurement than the first one was. ============================================================================== RUN better-auth--claude-opus-5--v7-b--2026-09-06 What Claude Opus 5 gets right about better-auth — battery v7-b, tested 2026-09-06 URL: https://stalepriors.com/runs/better-auth--claude-opus-5--v7-b--2026-09-06 JSON: https://stalepriors.com/data/better-auth/opus-5-v7-b.json Library: better-auth 1.7.3 (npm), verified 2026-09-06 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 1.2.0 (2025-03-01), lag ~26 months Oldest release it could not place: 1.3.0 (2025-07-19) Battery: better-auth/v7-b, 7 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The blind twin, and it disagreed with `v7-a` on the one quantity the battery was built around. From the identical stored prompt it answered task 1 "Yes" and task 5 `updatedAt` - the coherent pre-1.6.0 reading, held consistently across two tasks four apart - where its twin answered "no" and `createdAt`. So the freshness belief is **split within one subject**, and the split is not noise inside a draw: each arm is internally consistent and they disagree with each other. It reproduced its twin's two denials exactly (`resendStrategy`, `twoFactorPage`), which is what makes those a measurement. It charges nothing, by the duplicated-arm rule, and its task 1 miss has no sibling to point at because the charging arm passed. --- non-finding [miss] session.freshAge (measured from session.createdAt): Task 1, the primary probe, and the twin split this battery's P7 predicted would not happen. Where `v7-a` answered "no", this draw answered "Yes" - the pre-1.6.0 semantics - and stayed consistent with itself at task 5, naming `updatedAt` as the anchor and spelling out the reasoning that is exactly right for 1.5.0 and exactly wrong for 1.6.0: "The check compares `Date.now()` against the session's `updatedAt`, falling back to `createdAt` only when `updatedAt` is null. That is why task 1 passes: the row was touched two minutes ago, so it reads as fresh even though the session is 30 hours old." Executed against the installed packages, that request passes at 1.5.0 and is rejected with `FORBIDDEN / SESSION_NOT_FRESH` at 1.6.0 and 1.7.3. --- non-finding [miss] emailOTP({ resendStrategy }): Task 2. "No", with the absence claimed explicitly - "There is no `reuseOTP` / `allowResend`-style flag in the email-OTP plugin as far as I know" - and the same `generateOTP` cache its twin shipped, down to the in-process `Map` and the same two hedges (whether `generateOTP` may return a promise, and the verification row's identifier format). Both twins missed `resendStrategy` from the identical prompt, which makes this belief a measurement rather than a coin. --- non-finding [miss] twoFactorClient({ twoFactorPage }): Task 6. "No", and the reason given is a design argument rather than a memory: "It takes a callback, `onTwoFactorRedirect`, not a path string. The plugin can't do the navigation itself because it has no idea what router you're on, so there's nothing for a string to hook into." The shipped 1.6.0 client does exactly what the draw says it cannot: `window.location.href = options.twoFactorPage`, guarded by an `isSafeUrlScheme` check. It then hedged in the right direction without acting on it - "note that some early versions used a different name for this option". --- non-finding [correct] session.freshAge anchor option (does not exist): Task 3, the control. "No", correctly, and put more sharply than any other arm: "There's no 'measure freshness from createdAt vs updatedAt' switch. The freshness check reads one timestamp and that's it, so the security team and the product team are arguing about something the config surface does not expose." Note that it then offered `session.disableSessionRefresh: true` at "moderate, not full" confidence as a way to pin `updatedAt` near `createdAt` - a second name this session did not verify, flagged as an open question rather than scored. --- non-finding [correct] customSession: Task 4, the floor probe. `customSession` plus `customSessionClient()`, with the ordering constraint most draws did not mention ("`customSession` must be the last plugin in the array"). Passed. --- non-finding [context] session.freshAge (measured from session.createdAt): Task 5(a) carried the battery's most interesting piece of self-diagnosis, from the arm that got it wrong: "I want to flag a real discrepancy here: the documentation has described `freshAge` in creation terms ('fresh if the session was created within...'), while the implementation I remember uses `updatedAt`. Combined with `updateAge` sliding the session forward, that means an indefinitely active session can stay 'fresh' forever. If your threat model is 'prove you're still at the keyboard for sensitive actions', this is not the control you think it is." The draw had detected the exact discrepancy the vendor closed at 1.6.0, correctly identified which side the docs were on, and resolved it toward the stale implementation - and then recommended the reader check the source for their version, which would have corrected it. --- non-finding [context] stateless sessions: Task 7, the attribution anchor: the same denial as its twin, in the same shape. "I don't believe this release exists, and I'd push back on the premise rather than name a version... So if someone told you 'better-auth added stateless sessions in version X,' I'd want to see the changelog entry." Stateless session management shipped at 1.4.0. Belief data, never scored. ============================================================================== RUN better-auth--claude-opus-5--v1--2026-09-01 What Claude Opus 5 gets right about better-auth — battery v1, tested 2026-09-01 URL: https://stalepriors.com/runs/better-auth--claude-opus-5--v1--2026-09-01 JSON: https://stalepriors.com/data/better-auth/opus-5.json Library: better-auth 1.7.2 (npm), verified 2026-09-01 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 1.3.0 (2025-07-19), lag ~10 months Oldest release it could not place: 1.4.0 (2025-11-22) Battery: better-auth/v1, 12 tasks, tool uses during test: 0 Tested: 2026-09-01 Summary: Control arm of the milestone experiment, and the prediction held. Opus 5's describable boundary on better-auth is 1.3.0 (2025-07-19), an ordinary minor -- not the 1.0.0 milestone -- as the pre-registration required. Zero chargeable findings: it is the only subject to state the sign-in response shape exactly, the only one to surface that the bearer plugin accepts unsigned tokens by default, and it routed around the 1.4.0 stateless-session capability without denying it exists, which under the additive-API rule is an imprecision rather than a finding. It records the same attribution drift as the other two subjects, placing at least two features it genuinely knows in a release that did not contain them. Ten-month gap between its stated 2026-05 cutoff and its describable boundary here, which the subject itself predicted before being asked. --- non-finding [correct] signIn.email response: Task 1 -- gave the sign-in response shape exactly right, including the part the other two subjects got wrong: "It does not give you back a full `session` object -- that was true in very early versions but was removed." Verified against the shipped 1.7.2 package, which returns `{ redirect, token, url, user }`. --- non-finding [correct] bearer plugin: Task 5 -- correctly stated that `requireSignature: false` means a raw token from the database is accepted, and that this is a deliberate choice the caller must make. The only subject to surface the unsigned-token default. --- non-finding [correct] SSO plugin — SAML 2.0: Task 9 -- correct on SAML in the SSO plugin, correctly attributed to 1.3.0, and correctly flagged the plugin's move to a separate `@better-auth/sso` package. --- non-finding [correct]: Tasks 2, 3, 4, 6, 7, 8, 12 -- hooks, `oidcProvider`, SSO with `organizationProvisioning`, `admin.stopImpersonating`, `apiKey`, organization teams and `customSession` all correct and unusually detailed. --- non-finding [imprecision] stateless session management: Task 11 -- gave `secondaryStorage` with `storeSessionInDatabase: false` as the answer for keeping sessions out of the database, and said that for genuinely stateless verification "better-auth's own session is still the source of truth for minting; JWT is for service-to-service, not a replacement for the session store". It never mentions that omitting both `database` and `secondaryStorage` makes the cookie the session record. --- non-finding [context]: ATTRIBUTION DRIFT -- placed the device-authorization plugin and `lastLoginMethod` in 1.3.0. Neither appears in the 1.3.0 release notes; both are named in 1.4.0's, and there only in fix and improvement bullets, so their introducing release was not established. Also placed the MCP plugin in the "1.2/1.3 era" at stated ~75% confidence. --- non-finding [context]: Question (c) -- volunteered the limitation the Index's boundary metric depends on: "a late cutoff doesn't mean uniform coverage of every fast-moving npm package right up to it", and put its own effective boundary for this library nine months before its stated cutoff. ============================================================================== RUN better-auth--claude-sonnet-5--v2--2026-09-02 What Claude Sonnet 5 gets wrong about better-auth — battery v2, tested 2026-09-02 URL: https://stalepriors.com/runs/better-auth--claude-sonnet-5--v2--2026-09-02 JSON: https://stalepriors.com/data/better-auth/sonnet-5-v2.json Library: better-auth 1.7.2 (npm), verified 2026-09-02 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01 Version attribution stops at: 1.0.0 (2024-11-23), lag ~13 months Oldest release it could not place: 1.1.0 (2024-12-20) Battery: better-auth/v2, 5 tasks, tool uses during test: 0 Tested: 2026-09-02 Summary: The below-floor control arm, and it did half its job and failed the other half informatively. Claude Sonnet 5's attribution boundary sits at better-auth 1.0.0, four minors below the 1.3.8 surface this battery tests, so it is uninformative for the hypothesis by design — and it duly could not place either SAML SSO or device authorization to any release. Its job was to prove the capability probes can fail. On task 3 they can: it denied database-less sessions like every other draw, which is charged here as an S2. On tasks 1 and 2 they did not — it named both 1.3.8 plugins from four minors below, while volunteering that it might be reconstructing `deviceAuthorization` from generic RFC 8628 knowledge rather than remembering better-auth. That is the control catching a weakness in the instrument, which is what a control is for. --- F1 [S2 silently-wrong] Denies that better-auth can run with zero database, offering cookieCache and Redis instead API: stateless / database-less sessions Changed in better-auth 1.4.0 (2025-11-22), kind: added Chargeability: 1.4.0 (2025-11-22) precedes this subject's stated cutoff of 2026-01 by two months. Charged despite this subject's attribution boundary sitting far lower (1.0.0), because the fairness rule is about the release date against the stated cutoff, not about where the subject's recall happens to end. Model belief: "Honest caveat: I don't believe better-auth has a mode with zero database — it still needs persistent storage for the user/account tables. cookieCache avoids per-request DB reads for session checks; secondaryStorage lets you keep sessions in Redis/KV instead of your SQL database." Wrong: export const auth = betterAuth({ session: { cookieCache: { enabled: true, maxAge: 5 * 60 } }, secondaryStorage: { get: (key) => redis.get(key), set: (key, value, ttl) => redis.set(key, value, { EX: ttl }), delete: (key) => redis.del(key), }, }) Correct: // stateless since 1.4.0: omit BOTH `database` and `secondaryStorage` export const auth = betterAuth({ emailAndPassword: { enabled: true }, }) Impact: Same cost as on the other two subjects: a Redis stood up that the library has not required for session storage since 2025-11-22, or a database-less deployment target abandoned. Scope note: Verified at both ends per fact LF1. Not verified: that the feature is unchanged between 1.4.0 and 1.7.2. Source: https://github.com/better-auth/better-auth/releases/tag/v1.4.0 (2025-11-22) — "Stateless session management" Source: https://registry.npmjs.org/better-auth/-/better-auth-1.7.2.tgz (2026-08-26) — "the only place the session lives and therefore the authority itself (`false`, for stateless / DB-less deployments)" --- non-finding [context]: THE BELOW-FLOOR CONTROL ARM, and it returned UNINFORMATIVE for the hypothesis exactly as pre-registered — this subject's attribution boundary is 1.0.0, four minors below the 1.3.8 surface under test. It failed the internal control twice over: SAML enterprise SSO was "cannot place precisely ... this is my weakest answer of the four; treat it as closer to 'I have a hunch' than 'I recall this'", and device authorization was likewise "cannot place precisely". Under the rule fixed before the run, a subject that cannot place a genuine minor has attribution too noisy to read, so H is untestable here. Recorded as designed, not as a result. --- non-finding [context] deviceAuthorization(): THE CONTROL DID NOT DO ITS JOB ON TASKS 1 AND 2, AND THAT IS THE MOST USEFUL THING IT PRODUCED. Its pre-registered role was to show that the capability probes can elicit "the library does not offer this" from a subject whose boundary is far below the feature. It did not: it named `deviceAuthorization()` and `lastLoginMethod()` with their client plugins, four minors below where they shipped, while volunteering the reason to distrust that — "I can't fully rule out that I'm pattern-matching from generic OAuth Device Authorization Grant (RFC 8628) knowledge rather than a specific memory of better-auth shipping it." So task 1 probes a guessable name and cannot separate knowing from guessing, which is the failure mode that cost `langchain/v2` a whole battery. Task 2 is the sounder probe of the pair: `lastLoginMethod` is not an RFC name and nothing in the task suggests a library would ship it. The control succeeded in its other half — all five draws denied the task 3 capability, so the battery can elicit a denial and the charged findings are real. --- non-finding [correct] customSession(): Task 5, the floor probe, passed: `customSession()` with `customSessionClient()`, which the subject volunteered was "the one I recall with the most confidence of the five". The 1.0.0 floor is confirmed, so the low boundary recorded here is a measurement of this subject and not the battery probing beneath it. --- non-finding [correct] additional user fields in the sign-in response: Task 4 passed outright, with less hedging than either Opus or Fable draw: `const plan = data.user.plan` off the sign-in response, glossed "additionalFields ride along on the user object ... it's just the DB row serialized". Correct since 1.4.2 (2025-11-25). Five of five draws wrote this correctly, which falsifies the second half of pre-registered prediction P3. --- non-finding [context] oidcProvider(): A grading call recorded rather than buried. This draw's (c) named `oidcProvider`, `genericOAuth` and `multiSession` in a "~1.1 (rough guess, late 2024)" bucket, and the OIDC Provider plugin genuinely is 1.1.0 — which taken alone would move `knowledge_stops_at_version` up from the 1.0.0 that `better-auth/v1` recorded. It is not credited, because the same answer disclaims the entire mapping: "I don't have a reliable version-by-version changelog for this library in memory ... What follows is a fuzzy, low-confidence reconstruction of eras, not a verified list ... everything past 'very early 1.x' is already in 'version number with no reliable content attached' territory for me." Under the standing rule that a self-report is graded together with behaviour, an era-sketch its author explicitly refuses to stand behind is not a correct attribution. The bracket is therefore held at 1.0.0 / 1.1.0, matching `v1`. ============================================================================== RUN better-auth--claude-sonnet-5--v3-c--2026-09-03 What Claude Sonnet 5 gets right about better-auth — battery v3-c, tested 2026-09-03 URL: https://stalepriors.com/runs/better-auth--claude-sonnet-5--v3-c--2026-09-03 JSON: https://stalepriors.com/data/better-auth/sonnet-5-v3-c.json Library: better-auth 1.7.2 (npm), verified 2026-09-03 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01 Version attribution stops at: 1.2.0 (2025-03-01), lag ~18 months Oldest release it could not place: 1.3.0 (2025-07-19) Battery: better-auth/v3-c, 4 tasks, tool uses during test: 0 Tested: 2026-09-03 Summary: A below-floor control that did its job and then produced the battery's one genuinely new lead. It failed both capability probes, which is what establishes that neither is derivable from the surrounding API - without that, the two Opus denials could not be read as beliefs. It passed the floor probe and it did not invent the non-existent session-side option. The lead: asked about hashing verification identifiers, it did not hedge but asserted the capability does not exist in any version, and the per-plugin options it denied shipped in 1.3.0, eighteen months below its own stated cutoff. That is inside its window and it is not charged here, because the pre-registration made this arm a control and an arm cannot be re-designated once its results are read. --- non-finding [miss] baseURL as a dynamic multi-host config: Task 1, the control result the battery needed: "(i) No. As far as I know, `baseURL` in `betterAuth({...})` is a plain string, not a function or a per-request resolver. It's read once at config time." 1.5.0 (2026-03-01) is two months ABOVE this subject's stated cutoff, so the failure is expected and carries no information about staleness. Its value is the one thing a control is for: the probe is not derivable from the surrounding API, so the identical failure on the two Opus draws reads as a belief about the option rather than as an unguessable name. Prediction P2 confirmed on this arm. --- non-finding [miss] verification.storeIdentifier: THE ONLY RESULT IN THIS BATTERY THAT POINTS SOMEWHERE NEW, and it is deliberately not charged. On task 3 this draw said "(i) No, to my knowledge there's no built-in toggle for this either", and in (d)(ii) went further than any other draw: "I don't believe this exists in the library at any version. Not \"cannot place\" - I'm saying it doesn't exist, based on the absence of any recollection of such a flag despite reasonable familiarity with the verification-plugin surface (magic link, email OTP)." It then wrote a `databaseHooks.verification.create.before` transform and correctly identified, itself, that the transform breaks the read path. `magicLink({ storeToken })` and `emailOTP({ storeOTP })` shipped in **1.3.0 (2025-07-19)**, eighteen months BELOW this subject's stated cutoff - so this is a confident denial of a capability well inside its own window, not a staleness result. Two of the three other draws named those options correctly. --- non-finding [correct] session token hashing at rest: Task 2, the sibling control: "(i) No. I'm not aware of a config flag that hashes the session token before the row is written to `session`." Correct at every release; nothing invented. P3 holds on this arm. --- non-finding [correct] customSession(): Task 4, the floor probe, passed: "This one I'm fairly confident about - the `customSession` plugin exists specifically for this", with the correct plugin wiring. The run is therefore a measurement rather than a probe below the subject's knowledge. --- non-finding [context]: Two dating errors in the belief data, recorded because the Index tracks attribution separately from capability. This draw placed 1.0 at "around September 2024" (actual: 2024-11-23) and, in (d)(iv), stated "I don't believe better-auth's `sso` plugin supports SAML. My recollection is it's OIDC/generic-OAuth2 only." Fact LF4 records SAML in the SSO plugin at 1.3.0 (2025-07-19), inside this subject's window. Its (d)(iii) denial of database-less sessions matches every other draw and matches fact LF1 at 1.4.0. ============================================================================== RUN better-auth--claude-sonnet-5--v4-a--2026-09-03 What Claude Sonnet 5 gets right about better-auth — battery v4-a, tested 2026-09-03 URL: https://stalepriors.com/runs/better-auth--claude-sonnet-5--v4-a--2026-09-03 JSON: https://stalepriors.com/data/better-auth/sonnet-5-v4-a.json Library: better-auth 1.7.2 (npm), verified 2026-09-03 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff not stated Version attribution stops at: 1.1.0 (2024-12-20), lag ~null months Oldest release it could not place: 1.2.0 (2025-03-01) Battery: better-auth/v4-a, 3 tasks, tool uses during test: 0 Tested: 2026-09-03 Summary: Claude Sonnet 5, the arm this battery was built to charge, did not fail the probe. Asked about the plugins rather than about the `verification` table, it named `storeOTP` and `storeToken` with their correct defaults and value unions — the same subject that, under `better-auth/v3`'s table framing, asserted the capability does not exist at any version. It also correctly denied the same-scheme sibling on `phoneNumber`, so the pass reads as discrimination rather than as running the naming scheme. Nothing is charged. Two things this draw does establish: the framing was load-bearing, and Sonnet 5's belief about this surface is unstable — its blind twin, from the identical prompt, denied both options and reached for a database hook. The twins also disagree about their own cutoff, which is the second battery in which that has happened. --- non-finding [correct] magicLink storeToken / emailOTP storeOTP: Task 2, the target probe, and the prediction the battery was built on is FALSIFIED here. Asked verdict-first whether the magic-link and email-OTP plugins each take a storage option, this draw answered yes to both and named them: "email-OTP plugin: Yes. I recall an option — I believe named `storeOTP` — that accepts something like `"plain"` (default) / `"hashed"` / a custom hasher function" and "magic-link plugin: Yes, with lower confidence than the OTP one. I recall a parallel option, plausibly named `storeToken`, with the same shape". Both names, both defaults and both value unions are correct against the shipped 1.3.0 package. The same subject, under `better-auth/v3`'s verification-table framing, asserted the capability "doesn't exist in the library at any version". Per the battery's pre-registered scoring rule the pass is recorded as "did not deny it" rather than as knowledge — but this draw also correctly denied the same-scheme sibling on task 1, which is the discrimination pattern that reads as recall rather than as running the naming scheme. --- non-finding [correct] phoneNumber storeOTP (does not exist): Task 1, the same-scheme sibling control: correct denial, no invention. "No — not to my knowledge. I don't recall the `phoneNumber` plugin exposing anything like a `storeOTP` / `hashOTP` option the way I believe some of the other OTP-adjacent plugins do." `phoneNumber({ storeOTP })` is a TS2353 error at 1.3.0, 1.5.0 and 1.7.2. The draw named the hypothetical in the course of ruling it out, which JOURNAL/044 scores as a correct denial rather than an invention. --- non-finding [correct] databaseHooks.verification.create.before: Task 1 workaround reasoning, and it is right about the mechanism: it refused to hash via `databaseHooks.verification.create.before` alone, on the ground that the plugin's verify path compares the stored value against the submitted code, so hashing one side breaks sign-in. That is exactly the failure mode, and it is why the shipped per-plugin options own both halves of the comparison. --- non-finding [imprecision] magicLink storeToken custom-hasher form: Wrote the magic-link custom hasher as `storeToken: { type: "custom", hash }`. The shipped discriminant is `"custom-hasher"`, so the literal as written does not type-check. It was offered as an inline alternative in a comment, not as the primary answer, and the primary answer (`storeToken: "hashed"`) is correct. --- non-finding [miss] twoFactor otpOptions.storeOTP: Direct question (d)(iii) denied that the two-factor plugin has an at-rest hashing option for its OTP: "Cannot place, and I'm genuinely unsure this capability even exists in the form asked... If forced to guess whether it exists at all, I'd lean toward 'no, not as a dedicated option'." It does. `twoFactor({ otpOptions: { storeOTP } })` is present in `better-auth@1.3.0` and type-checks under `tsc --strict` at 1.3.0, 1.5.0 and 1.7.2 — it is the third of the four plugins that gained the option in that release. This is a real gap inside the fairness window and it is not charged, because the battery's pre-registration states that the direct questions are belief data and are never scored as findings. --- non-finding [miss] magicLink storeToken / emailOTP storeOTP: Version attribution, direct question (d)(i)/(ii): placed the per-plugin hashing options in the 1.2.x line — "My best estimate is this landed in the 1.2.x line (a minor or patch within it)" for the email-OTP option, and "cannot place" for magic-link. They shipped in **1.3.0** (2025-07-19). `storeOTP` and `storeToken` appear nowhere in the published `dist` of `better-auth@1.2.7` or `better-auth@1.2.12` — 1.2.12 being the last stable 1.2.x — and appear in five files of 1.3.0's. The subject demonstrated it holds the capability and then dated it one minor low, which is the case HARNESS.md's attribution rule says is readable. --- non-finding [context] magicLink storeToken / emailOTP storeOTP: The blind twins disagree on the target, and the disagreement runs the wrong way for the finding count. This charging draw named both options; the non-charging twin `better-auth--claude-sonnet-5--v4-b--2026-09-03` denied both from the same stored prompt and shipped a `databaseHooks` workaround. Sonnet 5 has now been drawn on this surface three times — `v3-c` denied it (control arm), `v4-a` named it, `v4-b` denied it (non-charging twin) — and BOTH denials landed in arms the scoring rules forbid from charging. That is the JOURNAL/029 undercount pattern, previously seen on the zod tuple surface, reproduced at a second library. --- non-finding [context]: The twins also disagree about their own training cutoff, from one stored prompt. This draw refused to name a month ("sometime in 2025"); `v4-b` accepted the environment-reported value ("Per the environment context I'm given, my cutoff is stated as January 2026"). That is the `zod/v4` instability (JOURNAL/031) reproduced at a second library and on a different subject pair, and it is now measured rather than assumed: the Sonnet 5 cutoff self-report is unstable across blind twins two batteries out of two. --- non-finding [correct] customSession: Floor probe (task 3) passed: named the `customSession` plugin, spread `user` and `session` back out of the callback, and added the companion `customSessionClient` on the client for type inference. The run is readable. ============================================================================== RUN better-auth--claude-sonnet-5--v4-b--2026-09-03 What Claude Sonnet 5 gets right about better-auth — battery v4-b, tested 2026-09-03 URL: https://stalepriors.com/runs/better-auth--claude-sonnet-5--v4-b--2026-09-03 JSON: https://stalepriors.com/data/better-auth/sonnet-5-v4-b.json Library: better-auth 1.7.2 (npm), verified 2026-09-03 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01 Version attribution stops at: 1.0.0 (2024-11-23), lag ~13 months Oldest release it could not place: 1.1.0 (2024-12-20) Battery: better-auth/v4-b, 3 tasks, tool uses during test: 0 Tested: 2026-09-03 Summary: The non-charging twin denied what its charging twin named. Asked whether the magic-link and email-OTP plugins take a storage option, this draw said no to both, would not invent a name, and shipped the designed wrong answer: hash on write through a database hook and hand-roll the verify. Both options shipped in 1.3.0, eighteen months below the cutoff this draw stated. It is not charged, because it is the `-b` draw of a duplicated arm. That makes two Sonnet 5 denials of this surface across three draws, both in arms the rules forbid from charging — the same shape that has kept the zod tuple miss uncharged for four sessions, now at a second library. It correctly denied the non-existent `phoneNumber` sibling and correctly described the deterministic-hash lookup it did not believe the library implements. --- non-finding [miss] magicLink storeToken / emailOTP storeOTP: Task 2, the target probe, reproduced the failure the battery exists to charge — in the arm that cannot charge it. Asked verdict-first, it answered "No, for both, to the best of my recollection. I do not have confident memory of a `magicLink` or `emailOTP` plugin option like `storeToken: "hashed"` or `storeOTP: "hashed"`... I can't name the option, and I won't invent one." Both options shipped in **1.3.0 (2025-07-19)**, eighteen months below this draw's own stated cutoff of 2026-01. It then shipped the designed wrong answer — hashing via `databaseHooks.verification.create.before` plus a hand-rolled verify — and repeated the denial in direct questions (d)(i) and (d)(ii) ("I don't believe this exists"). Verified by execution: at 1.3.0, `emailOTP({ storeOTP: "hashed" })` puts a digest in the row where the default puts `797478:0`, and sign-in with the raw code still succeeds. --- non-finding [correct] phoneNumber storeOTP (does not exist): Task 1, the same-scheme sibling control: correct denial, no invention, and unusually specific. It listed the plugin's real option surface — "`sendOTP`, `otpLength`, `expiresIn`, `allowedAttempts`, `signUpOnVerification` — none of which govern hashing at rest" — every one of which is a genuine member of `PhoneNumberOptions` at 1.7.2. --- non-finding [correct] magicLink storeToken / emailOTP storeOTP: Task 2(iii): although it denied the options exist, it explained the deterministic-hash lookup mechanism correctly — the identifier drives the lookup for the emailed code, the hashed token drives it for the magic link, and a per-row random salt would break both. That is exactly how the shipped options behave. Right about the mechanism, wrong about whether the library implements it. --- non-finding [miss] twoFactor otpOptions.storeOTP: Direct question (d)(iii) denied that the two-factor plugin has an at-rest hashing option for its OTP: "Cannot place. I have a vague, unreliable half-memory that sensitive secrets in `twoFactor` might be encrypted at rest using the app's core secret, but that's encryption of a TOTP secret, not hashing of a one-time code." It does. `twoFactor({ otpOptions: { storeOTP } })` is present in `better-auth@1.3.0` and type-checks under `tsc --strict` at 1.3.0, 1.5.0 and 1.7.2 — it is the third of the four plugins that gained the option in that release. This is a real gap inside the fairness window and it is not charged, because the battery's pre-registration states that the direct questions are belief data and are never scored as findings. --- non-finding [context]: Blind-twin disagreement on the target, reported on both runs and not resolved: this draw denied both options; `better-auth--claude-sonnet-5--v4-a--2026-09-03` named both from the identical stored prompt. The twins also disagree about their own training cutoff (this draw accepted 2026-01, `v4-a` declined to name a month). The `-b`-holds-the-better-answer pattern that HARNESS.md has tracked since JOURNAL/030 does NOT hold here — `-a` held the better answer — so that streak stands at four of seven. --- non-finding [correct] customSession: Floor probe (task 3) passed: named the `customSession` plugin, spread `user` and `session` back out of the callback, and added the companion `customSessionClient` on the client for type inference. The run is readable. ============================================================================== RUN better-auth--claude-sonnet-5--v5-c--2026-09-03 What Claude Sonnet 5 gets right about better-auth — battery v5-c, tested 2026-09-03 URL: https://stalepriors.com/runs/better-auth--claude-sonnet-5--v5-c--2026-09-03 JSON: https://stalepriors.com/data/better-auth/sonnet-5-v5-c.json Library: better-auth 1.7.2 (npm), verified 2026-09-03 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01 Version attribution stops at: 1.1.0 (2024-12-01), lag ~13 months Oldest release it could not place: 1.2.0 (2025-03-01) Battery: better-auth/v5-c, 3 tasks, tool uses during test: 0 Tested: 2026-09-03 Summary: The control arm that did not hold, and it is this battery's most consequential result. Pre-registered as a control because both of this subject's `better-auth/v4` draws denied the phone-number storage option correctly, it inverted: yes, `storeOTP`, shipped in the config as the fix. P2 falsified. That does not touch F1 - the charge on `v5-a` rests on non-compiling code, not on a contrast between arms - but it does change what the finding is about. Across the two batteries this belief is Opus 5 four draws for four, and Sonnet 5 one for three: the invention is not the property of a single subject that the pre-registration assumed, and the wording between v4 and v5 is the variable that moved. This draw also denied a real option on the one-time-token plugin, which the other three test draws all named. --- non-finding [miss] phoneNumber({ storeOTP }): Task 1, the probe, and the result that falsifies this battery's P2. This subject denied the same option correctly on BOTH of its `better-auth/v4` draws, four hours earlier, from a prompt asking about the same plugin and the same surface. Here it answered "(i) Yes. My recollection is that the `phoneNumber` plugin does accept an option that changes what gets persisted for the OTP", named `storeOTP` at "medium-high confidence", and shipped it with the comment "<- don't persist the raw code". `phoneNumber({ storeOTP })` is a TS2353 at 1.3.0, 1.5.0 and 1.7.2; the draw's entire remaining configuration - `sendOTP`, `otpLength`, `expiresIn`, `allowedAttempts` - is real and compiles. --- non-finding [miss] oneTimeToken storeToken: Task 2, the exists-control: denied `oneTimeToken({ storeToken })` - "one-time-token plugin: no (low confidence guess, not a confirmed fact)" - and restated it in (d)(ii) as "I don't believe this exists as a named plugin option at all". It shipped in 1.3.0 and type-checks clean at 1.3.0, 1.5.0 and 1.7.2. The other three test draws all named it correctly, two of them with its exact `{ type: 'custom-hasher', hash }` form. --- non-finding [correct] twoFactor otpOptions.storeOTP: Task 2, the other half: named `twoFactor({ otpOptions: { storeOTP } })` correctly, at the correct nesting, though on stated "medium confidence" and by explicit analogy to the phone-number option it had just invented. The right answer reached through a wrong premise - the analogy runs from a plugin that does not have the option to one that does. --- non-finding [correct] customSession: Task 3, the floor probe: passed. `customSession` named as the mechanism, correctly distinguished from `session.additionalFields` as "a different tool for a different job" - persisted columns versus a computed wrapper - on stated medium confidence. ============================================================================== RUN better-auth--claude-sonnet-5--v7-e--2026-09-06 What Claude Sonnet 5 gets right about better-auth — battery v7-e, tested 2026-09-06 URL: https://stalepriors.com/runs/better-auth--claude-sonnet-5--v7-e--2026-09-06 JSON: https://stalepriors.com/data/better-auth/sonnet-5-v7-e.json Library: better-auth 1.7.3 (npm), verified 2026-09-06 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01, SELF-TEST Version attribution stops at: 1.0.0 (2024-11-23), lag ~14 months Oldest release it could not place: 1.1.0 (2024-12-20) Battery: better-auth/v7-e, 7 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: A below-floor control, pre-registered as one, and it did the job a control exists to do: it answered the battery's primary probe **correctly**, from three months below the release that made that answer correct. That discounts `v7-a`'s pass on the same question - a derivable outcome kills a pass, not a failure - and leaves the two Fable 5.1 failures charged. It missed both named 1.6.0 options, which is expected and barred from charging by the fairness rule, and it was the only draw of the six to deny `twoFactorPage` without also asserting it had once existed and been removed. Its boundary answer on this library is lower than this subject's previous readings: it could not name a describable release above the early 1.x line. --- non-finding [correct] session.freshAge (measured from session.createdAt): Task 1, and the result that matters most in this arm: "No" - the **current**, post-1.6.0 answer, from a subject whose stated cutoff is 2026-01 and whose measured boundary on this library sits far below the release that changed it. It gave the reasoning too: "the freshness check (`session.freshAge`, default `60 * 60 * 24`, i.e. 24 hours) is measured from when the session was created (`createdAt`), not from recent activity." Correct on the anchor, correct on the default, and correct on the outcome, for a behaviour introduced three months after its cutoff. Task 5(a) repeated it: `createdAt`. **This is a derivability result and it is what a below-floor control is for.** Per JOURNAL/031, a derivable outcome discounts a pass and does not excuse a failure - so `v7-a`'s pass on the same question cannot be read as recall, while `v7-c`'s and `v7-d`'s failures still charge. Two caveats keep it honest: task 1 is binary, so a control agreeing is also a coin landing; and this draw's own boundary answer places its describable knowledge at the early 1.x line, i.e. before the pre-1.6.0 `updatedAt` behaviour was well documented - it may never have held the belief it would have had to un-learn. --- non-finding [miss] emailOTP({ resendStrategy }): Task 2: "No", with the absence claimed ("I don't recall the `emailOTP` plugin shipping a built-in 'reuse the existing code on resend' option") and a cache wired into `sendVerificationOTP` rather than `generateOTP` - a variant that sends the cached code but leaves the database row holding the newly generated one, so verification would reject the code the user was sent. The draw flagged the risk itself: "I'm not confident enough in the plugin's internals to promise this exactly matches its OTP-storage semantics." Every one of the six draws missed this option. --- non-finding [miss] twoFactorClient({ twoFactorPage }): Task 6: "No" - "My recollection is the `twoFactorClient` plugin takes a callback (`onTwoFactorRedirect`)... not a plain path string it navigates to on its own." Correct for its own era, and it did **not** reproduce the history inversion that three of the other draws produced: it made no claim that a string option had ever existed and been removed. --- non-finding [correct] session.freshAge anchor option (does not exist): Task 3, the control: "No", and it named the current implementation as the reason - "my recollection is the check is hardcoded against `createdAt`". Task 5(b) `freshAge: 0`, correct. The control holds in this arm. --- non-finding [correct] customSession: Task 4, the floor probe: `customSession` with a computed `isPro` field. The shortest correct answer of the six and it passed, so this arm is informative as a control. --- non-finding [context] stateless sessions: Task 7: declined rather than denied - "I don't have reliable, specific knowledge of a release number here" - and correctly separated `cookieCache` (a cache over a store) from the `jwt` plugin, while noting it could not confirm either eliminates the session store. Stateless session management is 1.4.0, which is above this subject's describable boundary. Belief data, never scored. ============================================================================== RUN better-auth--claude-sonnet-5--v1--2026-09-01 What Claude Sonnet 5 gets wrong about better-auth — battery v1, tested 2026-09-01 URL: https://stalepriors.com/runs/better-auth--claude-sonnet-5--v1--2026-09-01 JSON: https://stalepriors.com/data/better-auth/sonnet-5.json Library: better-auth 1.7.2 (npm), verified 2026-09-01 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01 Version attribution stops at: 1.0.0 (2024-11-23), lag ~13 months Oldest release it could not place: 1.1.0 (2024-12-20) Battery: better-auth/v1, 12 tasks, tool uses during test: 0 Tested: 2026-09-01 Summary: Test arm of the milestone experiment, and the prediction held. Sonnet 5's describable boundary on better-auth is 1.0.0 (2024-11-23) -- the milestone -- with 1.1.0 (2024-12-20) undescribable, exactly as pre-registered in prompts/better-auth.md, despite 1.1.0 being the larger release by release-note volume and named API surface. Two chargeable findings, both S2: it denies that database-less sessions exist (added 1.4.0) and, asked explicitly for concerns, omits that the bearer plugin accepts unsigned tokens by default. The floor probe passed, so the low boundary is a measurement rather than an artefact. The run's most consequential observation is not a finding: the subject used features from 1.1.0, 1.2.0 and 1.3.0 correctly while insisting it cannot describe anything past 1.0.0 -- feature knowledge and version attribution come apart, and it is attribution that this dataset's boundary metric measures. --- F1 [S2 silently-wrong] States that better-auth cannot run without a database-backed session store, which stopped being true in 1.4.0 API: stateless session management Changed in better-auth 1.4.0 (2025-11-22), kind: added Chargeability: 1.4.0 shipped 2025-11-22, inside the subject's stated 2026-01 cutoff. Model belief: Asked what the library offers a team that wants the session cookie smaller and no session state in the database, the subject answered: "I do **not** believe better-auth has a fully DB-less 'sessions never touch a database' mode -- a database adapter is core to how it works. So the honest answer is: you can shrink/avoid the cookie payload and get stateless *verification* via JWT, but you likely can't eliminate DB-backed session records entirely." Correct: // no `database`, no `secondaryStorage` -> the signed cookie IS the session record export const auth = betterAuth({ emailAndPassword: { enabled: true }, }) Impact: The team is told the architecture they asked for is unavailable in the library they are already using, and is steered toward either keeping a database they wanted to drop or re-platforming onto pure JWTs. The subject flagged its own uncertainty, which is to its credit, but it stated the negative rather than leaving it open, and that is the sentence a reader acts on. Scope note: Charged as a stated impossibility, not as a missing mention, under the additive-API rule in prompts/valibot.md v1: a model that routes around an added capability has written working code and is only imprecise; a model that says the capability does not exist is a finding. Claude Opus 5 routed around it on the same task and is recorded as an imprecision, not a finding, for exactly this reason. Source: https://github.com/better-auth/better-auth/releases/tag/v1.4.0 (2025-11-22) — "Stateless session management" Source: https://registry.npmjs.org/better-auth/-/better-auth-1.7.2.tgz (2026-08-26) — "the only place the session lives and therefore the authority itself (`false`, for stateless / DB-less deployments)" --- F2 [S2 silently-wrong] Asked to flag anything to know before shipping bearer-token auth, says nothing about the plugin accepting unsigned tokens by default API: bearer plugin Changed in better-auth 1.1.0 (2024-12-20), kind: behavior-changed Chargeability: 1.1.0 shipped 2024-12-20, thirteen months inside the subject's stated 2026-01 cutoff. Model belief: The subject mounted `bearer()` with no options and, asked explicitly to "flag anything about this we should know before shipping it", listed five considerations -- keychain storage, loss of httpOnly, manual expiry/refresh, explicit sign-out, CORS -- and did not mention signatures, `requireSignature`, or that the default accepts an unsigned token. Its configuration is the one that accepts unsigned tokens. Wrong: export const auth = betterAuth({ plugins: [bearer()] }) Correct: export const auth = betterAuth({ plugins: [bearer({ requireSignature: true })] }) Impact: A public API mounted this way accepts a raw session token with no signature check. The subject's own security advice was otherwise sound, which makes the omission more likely to be trusted, not less. Scope note: This is an OMISSION, not a false statement, and it is charged only because prompts/better-auth.md v1 designated this exact behaviour as a finding before the run: "A model ... answering the explicit 'anything we should know' with nothing about the default, gives security advice with a consequence." Recorded against the pre-registered rule rather than re-argued after seeing the data, which is the point of pre-registering it. A reader who thinks an omission should not be chargeable can subtract this one finding; it does not affect the boundary measurement or the prediction. Source: https://github.com/better-auth/better-auth/releases/tag/v1.1.0 (2024-12-20) — "By default the bearer plugin now accepts unsigned tokens and provides an option to require signed tokens only." Source: https://registry.npmjs.org/better-auth/-/better-auth-1.7.2.tgz (2026-08-26) — "@default false" --- non-finding [correct] hooks: Task 2 -- correctly used the global `hooks.before` / `hooks.after` config with `createAuthMiddleware` to run logic around every auth request, and noted that throwing there short-circuits. --- non-finding [correct] oidcProvider: Task 3 -- correctly named the `oidcProvider` plugin and the standard endpoints it exposes. --- non-finding [correct] stopImpersonating: Task 6 -- correctly named `admin.stopImpersonating()` and the preserved-original-session behaviour. --- non-finding [correct] apiKey: Task 7 -- correctly named the `apiKey` plugin with per-key rate limiting, expiry and deletion. --- non-finding [correct] organization teams: Task 8 -- correctly used `organization({ teams: { enabled: true } })` with separately managed team membership. --- non-finding [correct] customSession: Task 12 -- correctly used the `customSession` plugin to add a computed field to every session read. --- non-finding [correct] signIn.email response: Task 1 -- wrote `data.user.email` off the sign-in response, which is correct against the shipped 1.7.2 package. --- non-finding [imprecision] SSO plugin — SAML 2.0: Task 9 -- on SAML, answered "plausibly yes, verify the exact shape against current docs" and declined to invent API names, explicitly labelling its guess as a guess. --- non-finding [imprecision] deviceAuthorization: Task 10 -- named a `deviceAuthorization` plugin implementing RFC 8628 with the correct conceptual flow, at low stated confidence. --- non-finding [context] signIn.email response: Question (d) -- said a successful sign-in returns the user object plus "session information", "moderately confident there's also a `session` object or `token`". There is no `session` object in the response. --- non-finding [context]: THE ATTRIBUTION SPLIT -- the most important observation in this run. The subject placed its describable boundary at 1.0.0 and said "past v1.0, my knowledge stops being version-indexed at all -- it's a bag of features with no reliable version labels attached". Its feature knowledge is demonstrably far past 1.0.0: in the tasks it correctly used the hooks API, `oidcProvider`, the SSO plugin (all 1.1.0), `apiKey` and organization teams (1.2.0), and named SAML (1.3.0) and a device-authorization plugin. ============================================================================== RUN langchain--claude-fable-5-1--v4--2026-09-06 What Claude Fable 5.1 gets right about langchain — battery v4, tested 2026-09-06 URL: https://stalepriors.com/runs/langchain--claude-fable-5-1--v4--2026-09-06 JSON: https://stalepriors.com/data/langchain/fable-5-1-v4.json Library: langchain 1.4.0 (pypi), verified 2026-09-06 Model: Claude Fable 5.1 (claude-fable-5-1), stated cutoff 2026-06, SELF-TEST Version attribution stops at: 1.0.0 (2025-10-17), lag ~8 months Oldest release it could not place: 1.1.0 (2025-11-24) Battery: langchain/v4, 2 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: A boundary ladder, not a battery: nothing charged, `elicits_code: false`, one arm. Claude Fable 5.1 describes langchain through the 1.0 GA (2025-10-17) and holds everything after it as a bare version number, eight months below its stated June 2026 cutoff. That is the same boundary Claude Opus 5 and Claude Sonnet 5 have shown across four earlier langchain batteries — three subjects, three stated cutoffs, one stopping point. Both ends of the ladder held: the poison rung (0.4.0, never published) was rejected with the correct reason, and 1.4.0 — published three days before this draw — was declined. One rung disagrees with its own explanation and the disagreement is published rather than resolved silently. --- non-finding [context] boundary: THE CELL THIS ARM WAS RUN TO FILL. Claude Fable 5.1 had never been drawn on langchain. Boundary: describes 1.0.0 (2025-10-17), holds 1.1.0 (2025-11-24) as a number only — the same boundary Claude Opus 5 and Claude Sonnet 5 landed on across langchain/v1, /v1r, /v2 and /v3, so three subjects with three different stated cutoffs (2026-01, 2026-05, 2026-06) all stop describing this library at the 1.0 GA. Consequence for the sweep: langchain 1.1.0, 1.2.0 and 1.3.0 are LIVE against this subject, and 1.3.0 (2026-05-12) carries no finding from any subject. --- non-finding [context] rung sort vs body — 1.2.0: THE GRADED WORD AND THE EXPLANATION DISAGREE ON ONE RUNG, AND THE INSTRUMENT SAYS WHICH ONE COUNTS. The sort answered DESCRIBE for 1.2.0; the explanation two lines later said "1.1.0 and 1.2.0 I believe exist (late 2025 / early 2026) but I can only vaguely gesture at their contents, so NAME". This is HARNESS § *The hedge rule run backwards* in a second place — the single word is the graded verdict, so the graded sort reads 1.2.0 as DESCRIBE — but no content for 1.2.0 was ever produced, in the sort or in direct question (c), and (a) and (c) both place the last describable release at 1.0.0. The boundary is recorded at 1.0.0 on three agreeing statements with the fourth disclosed here. Under the alternate reading it is 1.2.0, which would make langchain 1.3.0 the only live release rather than three. --- non-finding [correct] poison rung 0.4.0, and the top rung 1.4.0: BOTH ENDS OF THE LADDER HELD. The poison rung was rejected with the right reason — "I do not believe a 0.4.0 ever shipped; the line went 0.3.x straight to 1.0" — which is exactly the published history. The top rung, 1.4.0, shipped 2026-09-03, three days before this draw and three months above the subject's stated cutoff, and was answered NO, with the honest gloss that "given the release cadence it may well exist by now". An arm that rejects a version that does not exist and declines one that exists but postdates it is an arm whose other version answers can be read. --- non-finding [correct] TASK 1 — create_agent: The demonstration task. `from langchain.agents import create_agent`, `init_chat_model`, `@tool`, `system_prompt=` with the note that pre-1.0 called it `prompt=`, `result["messages"][-1].content`, and middleware / checkpointer / response_format named as the extension points. That is the 1.0 GA surface, correctly attributed to 1.0 and correctly distinguished from the legacy `AgentExecutor` and `langgraph.prebuilt.create_react_agent` paths. --- non-finding [imprecision] 1.1.0 content, volunteered: Volunteered inside direct question (c) while declining to describe 1.1.0: "I believe it shipped around November 2025 and have a vague sense it touched model-capability profiles and middleware refinements." The date is right to the month (2025-11-24) and 'middleware refinements' is directionally right — langchain/v3 charged three subjects on 1.1.0 middleware defaults. Recorded as an imprecision rather than a finding because it is hedged as a vague sense and asserts nothing specific enough to be wrong; it is also the reason this arm's boundary is worth publishing with its disagreement visible rather than flattened. ============================================================================== RUN langchain--claude-fable-5-1--v5-a--2026-09-06 What Claude Fable 5.1 gets wrong about langchain — battery v5-a, tested 2026-09-06 URL: https://stalepriors.com/runs/langchain--claude-fable-5-1--v5-a--2026-09-06 JSON: https://stalepriors.com/data/langchain/fable-5-1-v5-a.json Library: langchain 1.4.0 (pypi), verified 2026-09-06 Model: Claude Fable 5.1 (claude-fable-5-1), stated cutoff 2026-06, SELF-TEST Version attribution stops at: 1.1.0 (2025-11-24), lag ~6 months Oldest release it could not place: 1.2.0 (2025-12-15) Battery: langchain/v5-a, 7 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The charging arm of the battery pre-registered against langchain 1.3.0. Two findings charged, both S3, both from the same four-month blind spot: `create_agent(transformers=...)` denied in the verdict (task 1) and again in the artefact (task 5), and `v2` named as the ceiling of `astream_events` (task 6) when 1.3.0 added `v3` for agents. The control sibling held, the floor passed, and the 11k-i deprecation pair came back four-for-four correct - including `extras`, a 1.2.0 addition inside this subject's blind window, which it described accurately at low confidence. The pre-registered P1 - that the named surface would draw more wrong answers than the small-answer-space one - is falsified by this arm and by the whole battery: both surfaces drew wrong answers from all five arms. --- F1 [S3 deprecated] Denies that `create_agent` can register stream transformers at all, then ships a consumer-side wrapper that rebuilds per-scope identity by hand - the verdict and the artefact wrong in the same direction, four tasks apart API: create_agent(transformers=...) Changed in langchain 1.3.0 (2026-05-12), kind: added Chargeability: 1.3.0 shipped 2026-05-12; this draw states a June 2026 cutoff, which is after it and not the same month, so the fairness rule and the same-month bar both clear. This is the pre-registered charging arm and tasks 1 and 5 are pre-registered probes on this surface. **One finding, two artefacts** (JOURNAL/062): the pre-registration graded the verdict (task 1) and the artefact (task 5) independently and this draw got both wrong in the same direction, which is one belief measured twice. Model belief: Task 1, single word on its own line: "No", glossed "I am not aware of a public 'stream transformer' / scope-aware factory API on the compiled graph at all; if one exists it is newer than what I can describe with confidence." Task 5, four tasks later: "The library, as I know it, gives me no way to register a transformer on the compiled graph, so here is what I would actually ship" - a `TransformingAgent` wrapper that calls `.astream(..., subgraphs=True)` and keys a `dict` of transformer instances off the namespace tuple to reconstruct the per-scope property the parameter already provides. Task 7, on which release added it: "I cannot name one. In every release I can describe ... there is no way to register your own stream transformers on the compiled agent." The subject also listed the parameters it believes `create_agent` takes - model, tools, system_prompt, middleware, response_format, state_schema, context_schema, checkpointer, store, interrupt_before/after, debug, name, cache - which is exactly the 1.2.18 signature with `transformers` missing. Wrong: # The artefact from task 5: the capability is denied, so per-scope identity is # reconstructed on the consumer side from subgraph namespace tuples. class TransformingAgent: async def astream(self, inputs, *, stream_mode="updates", config=None): per_scope: dict[tuple, MyTransformer] = {} async for ns, mode, chunk in self._graph.astream( inputs, config=config, stream_mode=[stream_mode], subgraphs=True ): if ns not in per_scope: per_scope[ns] = self._factory(ns) # invoked once per scope yield per_scope[ns](mode, chunk) Correct: # 1.3.0 and later: the factory is registered on the graph the agent compiles, # after the built-in ToolCallTransformer, and is invoked once per scope. agent = create_agent( model, tools, system_prompt="...", transformers=[MyTransformer], ) Impact: The wrapper runs, which is why this is S3 rather than S1: a reader who follows it ships working code and never learns the parameter exists. What it costs is the wrapper itself, and the guarantee that comes with the supported path - the agent registers `ToolCallTransformer` first and appends yours after it, so the built-in behaviour is kept rather than reimplemented. Source: https://files.pythonhosted.org/packages/7b/6f/b9a9721c27fbb6d29a6a7cd89d6a41eeffc7c79b49b9a5cf5beb1d60952d/langchain-1.3.0-py3-none-any.whl (2026-05-12) — "transformers: Optional sequence of scope-aware `StreamTransformer` factories to register on the compiled graph in addition to the agent defaults. Each factory is invoked per-scope (`factory(scope)`) so subgraph mini-muxes get fresh instances. Appended after the built-in `ToolCallTransformer`." Source: https://files.pythonhosted.org/packages/59/20/959f6098c79158afe5aedce7de05c3700f10d293890ef9e5dace6c3ad94b/langchain-1.2.18-py3-none-any.whl (2026-05-08) Source: https://docs.langchain.com/oss/python/releases/changelog (2026-05-12) — "This release adds support for version="v3" in stream_events / astream_events for langchain agents." --- F2 [S3 deprecated] Names `v2` as the highest event-stream protocol version an agent accepts, and describes it as the ceiling - v3 shipped for langchain agents in the same release as F1's parameter API: astream_events(version="v3") on a create_agent agent Changed in langchain 1.3.0 (2026-05-12), kind: added Chargeability: Same window and same licence as F1. Charged separately from F1 because it corrects a different published fact (LF38, not LF37) on a different API surface; JOURNAL/062's one-finding rule is about one belief measured by two tasks, not about two surfaces that happen to ship in one release. Model belief: Task 6(a), the string on its own line: "v2". Task 6(b): "`v2` gives a consistent, normalised event schema: `parent_ids` on every event, `on_chat_model_end`/`on_tool_end` outputs as the actual message/tool result rather than wrapped `LLMResult`-style objects, and consistent naming/ordering of nested events - `v1` had known inconsistencies that `v2` was introduced to fix, and it is the version the docs and the `create_agent` graph are exercised against." Every word of that is a correct description of the v1 -> v2 change; what is wrong is that it is offered as the top of the ladder. Wrong: async for event in agent.astream_events(inputs, version="v2"): ... Correct: async for event in agent.astream_events(inputs, version="v3"): ... # run.values / run.messages / run.lifecycle / run.subgraphs Impact: v1 and v2 are unchanged at 1.3.0, so the code the belief produces still runs; the cost is that the consumer hand-builds typed per-channel projections the protocol now supplies. Source: https://docs.langchain.com/oss/python/releases/changelog (2026-05-12) — "This release adds support for version="v3" in stream_events / astream_events for langchain agents." Source: https://docs.langchain.com/oss/python/releases/changelog (2026-05-12) — "Pass version="v3" to stream_events() / astream_events() for a content-block-centric protocol with typed, per-channel projections (run.values, run.messages, run.lifecycle, run.subgraphs) ... version="v1" and version="v2" are unchanged." Source: https://files.pythonhosted.org/packages/7b/6f/b9a9721c27fbb6d29a6a7cd89d6a41eeffc7c79b49b9a5cf5beb1d60952d/langchain-1.3.0-py3-none-any.whl (2026-05-12) — "transformers: Optional sequence of scope-aware `StreamTransformer` factories to register on the compiled graph in addition to the agent defaults. Each factory is invoked per-scope (`factory(scope)`) so subgraph mini-muxes get fresh instances. Appended after the built-in `ToolCallTransformer`." --- non-finding [correct] create_agent stream-mode parameter (does not exist): TASK 2, THE PRE-REGISTERED CONTROL SIBLING, AND IT HELD - WITH A WRINKLE WORTH RECORDING. "No", correctly: `create_agent` has no parameter that fixes the compiled agent's default streaming mode, at any release. The arm then volunteered a mechanism the spec had not anticipated - "Pregel exposes a `stream_mode` attribute that `.stream()/.astream()` fall back to when the caller passes none" - hedged as "an attribute, not a documented constructor option". That claim is TRUE: langgraph 1.2.11 `pregel/main.py` declares `stream_mode: StreamMode = "values"` on the Pregel class. So the sibling is half-real - absent as a `create_agent` parameter, present as an attribute on the object it returns - which is a better control than a wholly imaginary one, because the arm had a true thing to reach for and still answered the question that was asked. --- non-finding [correct] create_agent floor probe (1.0.0): TASK 3, THE FLOOR PROBE, PASSED. `from langchain.agents import create_agent`, tools as `@tool`-decorated functions, `system_prompt=`, `agent.invoke({"messages": [...]})`, final answer read off `result["messages"][-1].content`, and the correct note that `from langchain_core.tools import tool` and `langchain.tools` are the same decorator. Nothing in this arm's control answers is discounted. --- non-finding [correct] @tool(extras=...) / create_agent(middleware=...): TASK 4, BACKLOG 11k-i's DEPRECATION PAIR, AND ALL FOUR VERDICTS ARE RIGHT. `extras`: available yes, ever deprecated no - with the mechanism correct ("an `extras` field being added to `BaseTool` / the `@tool` decorator in the langchain-core 1.x line (late 2025) to carry provider-specific tool metadata ... through to the provider's tool schema") and flagged as low confidence. `middleware`: available yes, ever deprecated no, flagged high confidence. `extras` is LF26, added 1.2.0 (2025-12-15) - inside this subject's blind window and above its boundary - so this is a PASS on a live release, not a floor answer. **No fabricated deprecation appeared**: see the run summary. --- non-finding [correct] task 7 attribution: TASK 7 IS AN ABSTENTION, NOT AN INVENTION, AND IT IS RECORDED AS THE RIGHT BEHAVIOUR FOR A SUBJECT IN THIS POSITION. Asked which release first allowed registering stream transformers, the arm said "I cannot name one ... If a release added it, it is after the point where my knowledge of langchain's release contents ends, and I would be guessing at a version number." JOURNAL/046: a denial is not an invention. The denial of the capability is charged as F1; declining to date it is the one part of the answer that is well calibrated. --- non-finding [context] boundary: BOUNDARY SPREAD ON THIS PAIR IS NOW TWO MEASUREMENTS AND THEY DISAGREE BY ONE RELEASE. The `langchain/v4` ladder placed Claude Fable 5.1 at 1.0.0 (describes 1.0.0, holds 1.1.0 as a number). This draw describes 1.1.0 in detail - "additional middleware ... model-profile awareness for capabilities, and fixes around structured output in `create_agent`" - and names 1.2 as the first release it knows only as a number. So the readings are 1.0.0 / 1.1.0, and per BACKLOG 11k-l a future battery in that band must treat the floor as the HIGHEST reading, 1.1.0. Nothing charged here turns on it: 1.3.0 is above both, and LF26 (1.2.0) was answered correctly rather than charged. ============================================================================== RUN langchain--claude-fable-5-1--v5-b--2026-09-06 What Claude Fable 5.1 gets right about langchain — battery v5-b, tested 2026-09-06 URL: https://stalepriors.com/runs/langchain--claude-fable-5-1--v5-b--2026-09-06 JSON: https://stalepriors.com/data/langchain/fable-5-1-v5-b.json Library: langchain 1.4.0 (pypi), verified 2026-09-06 Model: Claude Fable 5.1 (claude-fable-5-1), stated cutoff 2026-06, SELF-TEST Version attribution stops at: 1.1.0 (2025-11-24), lag ~6 months Oldest release it could not place: 1.2.0 (2025-12-15) Battery: langchain/v5-b, 7 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The blind twin of `v5-a`, charging nothing by design. It reproduced both failures - the `transformers` denial in the verdict, the artefact and the attribution, and `v2` as the `astream_events` ceiling - and passed the control, the floor and all four 11k-i verdicts. The one place the pair differs is the shape of the workaround: `v5-a` wrapped the consumer side, `v5-b` wrote a middleware that emits custom events. Same belief, two artefacts. --- non-finding [miss] create_agent(transformers=...): THE SAME DENIAL AS THE SIBLING, IN BOTH PLACES. Task 1: "no". Task 5: "To my knowledge the library gives you no hook for registering a stream transformer on the graph that `create_agent` compiles", followed by a `StreamTransformMiddleware` that calls `get_stream_writer()` inside `wrap_model_call` and emits custom events - a different workaround from the sibling's, reaching the same place - plus the explicit statement that "if you genuinely need a factory invoked once per scope, that concept doesn't exist in LangGraph as I know it". Task 7: "I can't name one. As far as I know, no release of `langchain` has ever let you register your own stream transformers on a `create_agent` agent." The pair agrees on the verdict, the artefact and the attribution; the two workarounds differ, which is the instrument spread on this probe. --- non-finding [miss] astream_events(version="v3") on a create_agent agent: "v2", identically to the sibling, with a correct account of what v2 fixed in v1 and the additional claim that v1 was "removed in langchain-core 1.0". Both arms of the pair treat v2 as the ceiling. --- non-finding [correct] create_agent stream-mode parameter (does not exist): TASK 2, THE PRE-REGISTERED CONTROL SIBLING, AND IT HELD - WITH A WRINKLE WORTH RECORDING. "No", correctly: `create_agent` has no parameter that fixes the compiled agent's default streaming mode, at any release. The arm then volunteered a mechanism the spec had not anticipated - "Pregel exposes a `stream_mode` attribute that `.stream()/.astream()` fall back to when the caller passes none" - hedged as "an attribute, not a documented constructor option". That claim is TRUE: langgraph 1.2.11 `pregel/main.py` declares `stream_mode: StreamMode = "values"` on the Pregel class. So the sibling is half-real - absent as a `create_agent` parameter, present as an attribute on the object it returns - which is a better control than a wholly imaginary one, because the arm had a true thing to reach for and still answered the question that was asked. --- non-finding [correct] create_agent floor probe (1.0.0): TASK 3, THE FLOOR PROBE, PASSED, and with the sharpest gloss any arm gave: "`from langchain.agents import create_agent` (langchain 1.x; the legacy `initialize_agent`/`AgentExecutor` path now lives in `langchain-classic`)", plus the correct `{"messages": [HumanMessage(...)]}` invoke shape. --- non-finding [correct] @tool(extras=...) / create_agent(middleware=...): TASK 4: four verdicts, all right - yes / no / yes / no - with `extras` correctly placed in "the langchain-core 1.x line, roughly late 2025" and correctly described as carrying provider-specific tool parameters such as Anthropic `cache_control`. Like the sibling, this arm produced NO fabricated deprecation on either option. --- non-finding [context] boundary: The twin's boundary reading matches the sibling's exactly - describes 1.1.0, hazy on 1.2, and it names the first number-only release explicitly: "The first release I know only as a version number is 1.3.0." Both blind draws therefore place their own blind spot one release below the target of the battery, which is the cleanest statement of the window this corpus has. ============================================================================== RUN langchain--claude-fable-5-1--v6-e--2026-09-06 What Claude Fable 5.1 gets right about langchain — battery v6-e, tested 2026-09-06 URL: https://stalepriors.com/runs/langchain--claude-fable-5-1--v6-e--2026-09-06 JSON: https://stalepriors.com/data/langchain/fable-5-1-v6-e.json Library: langchain 1.4.0 (pypi), verified 2026-09-06 Model: Claude Fable 5.1 (claude-fable-5-1), stated cutoff 2026-06, SELF-TEST Version attribution stops at: 1.2.0 (2025-12-15), lag ~6 months Oldest release it could not place: 1.3.0 (2026-05-12) Battery: langchain/v6-e, 6 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: Pre-registered as a control and it behaved like the instrument it was chosen to be. `langchain/v5` handed this subject the name `extras` and it recognised it; this battery never says the word, and the subject produced it anyway — one-word verdict "Yes" on task 1 with `extras` named in the same breath, the correct **flat** shape shipped in task 4 and executed to confirm both provider fields reach the wire, and an attribution in task 6 that puts the field in the langchain-core 1.1/1.2 line in December 2025 against an artifact of 1.2.0 on 2025-12-12. Pre-registered prediction P2 holds: those `v5` answers were recall, not recognition. It also passed the floor cleanly, denied the sibling correctly while naming the only correct module path for `convert_to_anthropic_tool` in the whole battery, got both `response_format` verdicts right, and produced no fabricated deprecation in the one arm where the 11k-i probe was in its designed configuration. Charges nothing, by a designation fixed before the draws. --- non-finding [correct] @tool(extras={...}): **THE RESULT THE BATTERY WAS BUILT AROUND, AND PRE-REGISTERED PREDICTION P2 HOLDS: THE NAME CAME OUT UNPROMPTED.** `langchain/v5` handed this subject the string `extras` and it recognised it. This battery never says the word, and this draw produced it anyway. Task 1's one-word verdict: "Yes", immediately glossed "(I believe this is `extras` on `BaseTool` / the `@tool` decorator — see task 4. Moderate confidence.)" Task 4 shipped `@tool(extras={"cache_control": {"type": "ephemeral"}, "defer_loading": True})` — the **flat** shape, keyed by the provider's own field names, which is the component the pre-registration marked as *not* derivable. Task 4's MECHANISM line: "extras". Executed on 2026-09-06 against langchain-core 1.6.2 and langchain-anthropic 1.7.1: `convert_to_anthropic_tool` on that tool returns `{'name': 'search_docs_extras', 'input_schema': {...}, 'description': ..., 'cache_control': {'type': 'ephemeral'}, 'defer_loading': True}` — both fields on the wire. Run the same probe with the nested shape LF26 records as wrong, `extras={"anthropic": {...}}`, and both fields are dropped by the `_ANTHROPIC_EXTRA_FIELDS` filter; this draw did not write that. So `v5`'s two correct answers were recall, not recognition, on the one subject where the two could be told apart. --- non-finding [correct] the introducing release for tool extras (task 6 attribution): The most accurate attribution any subject has produced on this surface: "the field arrived in `langchain-core` 1.x, shortly after the 1.0 release (1.0.0 shipped mid-October 2025) — I place it in the langchain-core 1.1 / 1.2 line, roughly December 2025 to January 2026, with `langchain-anthropic` gaining the corresponding passthrough (for `cache_control`, `defer_loading`, and later `input_examples`) at about the same time, driven by Anthropic's tool-search / deferred-tools feature (November 2025). I'm not confident about the exact minor version." The artifact: langchain-core **1.2.0, 2025-12-12**; the `langchain` changelog entry **2025-12-15**. The right minor is inside the stated 1.1/1.2 range, the month is exact, and the three field names it lists are exactly three of the five in langchain-anthropic's `_ANTHROPIC_EXTRA_FIELDS`. Compare JOURNAL/028's minor-granularity rule: this is an attribution correct to the minor, stated with a hedge that names the right two candidates. --- non-finding [correct] a provider-schema-format parameter on @tool (the sibling that does not exist): Task 2(a): "No" — correct, and it named both nearby true things with their correct modules: `convert_to_openai_tool` and `convert_to_json_schema` in `langchain_core.utils.function_calling`, and `convert_to_anthropic_tool` in `langchain_anthropic.chat_models`. It is the only arm in the battery that placed the Anthropic converter in the right package (`v6-c` doubted it exists; `v6-d` imported it from `langchain_core`). No charge from this task in either direction. --- non-finding [correct] create_agent (the floor probe, langchain 1.0.0): Task 3: passed cleanly, and above the floor. `from langchain.agents import create_agent` with `system_prompt=`, `init_chat_model("anthropic:claude-sonnet-4-5")`, and the final answer read off `result["messages"][-1].text` — the `text` property on the message, which is verified present at langchain-core 1.6.2 and is the 1.x-era accessor rather than `.content`. --- non-finding [correct] @tool(response_format="content_and_artifact") — the supplied-name control half: Task 5(b): "Yes" available, "No" never deprecated — both correct, dated to langchain-core 0.2.x mid-2024 and correctly stated to be unchanged in 1.x. --- non-finding [context] the fabricated-deprecation probe (BACKLOG 11k-i, rebuilt): The one arm in the battery for which this probe was in its designed configuration — a **correct** name, **self-produced**, held at explicitly moderate confidence, which is the `better-auth/v7` shape minus nothing. Task 5(a): "Yes" available, "No" never deprecated, glossed "`extras` — as far as I know — is a recent addition (LangChain 1.x era) and has not been renamed or deprecated since; it was preceded not by another name but by the 'pass a raw dict' workaround. My confidence that `extras` is the exact name is moderate, not high; if it is wrong, the real name is something very close in spirit." **No fabricated deprecation**, and the uncertainty was discharged as a hedge on the *name* rather than as an invented removal. This is the strongest single piece of evidence for P3. --- non-finding [context] ChatAnthropic(betas=[...]) and the tool-search server tool: Unprompted and correct in outline: the draw noted that `defer_loading` "only makes sense alongside Anthropic's tool-search server tool, and both currently require the advanced-tool-use beta header", and shipped `ChatAnthropic(model=..., betas=["advanced-tool-use-2025-11-20"])` plus a `{"type": "tool_search_tool_regex_20251119", "name": "tool_search_tool_regex"}` entry in the bound tool list. `betas` is verified a real field on `ChatAnthropic` at langchain-anthropic 1.7.1. The two dated beta/tool identifiers were not verified this session and nothing here turns on them; they are recorded as unverified, exactly as BACKLOG 11k-j records an unverified name rather than citing it. ============================================================================== RUN langchain--claude-fable-5--v2--2026-09-01 What Claude Fable 5 gets wrong about langchain — battery v2, tested 2026-09-01 URL: https://stalepriors.com/runs/langchain--claude-fable-5--v2--2026-09-01 JSON: https://stalepriors.com/data/langchain/fable-5-v2.json Library: langchain 1.3.18 (pypi), verified 2026-09-01 Model: Claude Fable 5 (claude-fable-5), stated cutoff 2026-01 Battery: langchain/v2, 5 tasks, tool uses during test: 0 Tested: 2026-09-01 Summary: The strongest performance in this battery, and it proves nothing. Three of the four counted surfaces came back clean and unhedged — `model.profile` with `image_inputs`, a `SystemMessage` carrying `cache_control` as `system_prompt`, and `ModelRetryMiddleware` led with rather than offered as a fallback, every keyword valid against the shipped signature. On the fourth it named `extras` correctly and put it in the right place, then nested the values under a provider key the attribute does not use, so the flags reach the provider as an `anthropic` field it will ignore. It is also the only subject that dated anything correctly: it placed the retry middleware at 1.1, ~Dec 2025, against an actual 1.1.0 on 2025-11-24. Three of four meets the pre-registered pass threshold — and the outcome table fixed before the run voids that reading, because the control arm scored two of four. Recorded as UNINFORMATIVE about the hypothesis, which is what the pre-registration says to do. --- F1 [S2 silently-wrong] Names `extras` correctly but nests the values under a provider key the shipped attribute does not use API: tool extras Changed in langchain 1.2.0 (2025-12-15), kind: added Chargeability: langchain 1.2.0 shipped 2025-12-15 and langchain-core 1.2.0 on 2025-12-12, both inside the subject's stated 2026-01 window. Model belief: "My belief: recent langchain 1.x lets you attach provider-specific parameters to a tool itself via a per-provider `extras` mapping on the tool." Confidence stated as "low-to-moderate on `extras` being the exact attribute name and `@tool(extras=...)` the exact spelling". Wrong: @tool(extras={"anthropic": {"defer_loading": True}}) def giant_schema_tool(query: str, options: dict, filters: dict) -> str: ... @tool(extras={"anthropic": {"cache_control": {"type": "ephemeral"}}}) def hot_tool(x: str) -> str: ... Correct: @tool(extras={"defer_loading": True}) def giant_schema_tool(query: str, options: dict, filters: dict) -> str: ... @tool(extras={"cache_control": {"type": "ephemeral"}}) def hot_tool(x: str) -> str: ... Impact: `extras` is typed `dict[str, Any]`, so the provider-keyed dict is accepted and nothing raises. The provider then receives a tool carrying an `anthropic` field it does not understand instead of the `defer_loading` and `cache_control` fields it does, so neither deferral nor caching takes effect. The symptom is identical to Opus 5's invented attribute — a token bill that does not fall — reached by a much closer miss. Scope note: DEVIATION FROM THE PRE-REGISTRATION, disclosed. The battery fixed a severity ceiling of S3 for all v2 probes before the run, reasoning that failing to use an *addition* cannot break a build. That reasoning was wrong in a way the run exposed: it conflated "cannot break a build" (true) with "cannot be silently wrong" (false). This failure is silently-wrong — the code runs and the provider instruction is discarded — and the site renders severity and label as one four-point scale, so filing it S3 would publish the blurb "works today, on a path the library has deprecated", which is false about this finding. Scored S2 because publishing an accurate description outranks honouring a ceiling that was misdrawn. The bias risk is named rather than hidden: raising a severity after seeing the data flatters the Index numbers, and three findings in this battery move S3 -> S2 because of it. The standing rule is amended (BACKLOG.md) so future ceilings are set by failure mode, not by change kind. Not executed: established from the shipped package, not from a run. Graded a partial rather than a pass under the battery's pre-registered rule for "names the right API with a wrong signature", and shipped as a finding because the value shape is part of the calling convention and the wrong shape changes what reaches the provider. Source: https://docs.langchain.com/oss/python/releases/changelog (2025-12-15) — "Simplified support for provider-specific tool parameters and definitions via a new extras attribute on tools." Source: https://files.pythonhosted.org/packages/8e/25/f50dd65673c819aa33d3c34df58c115dbb6ec627d19f93e6e401dd0fc8d7/langchain_core-1.6.1-py3-none-any.whl (2026-08-27) — "@tool(extras={"defer_loading": True, "cache_control": {"type": "ephemeral"}}) def my_tool(x: str) -> str:" --- non-finding [correct] model profiles (.profile): P1 — `getattr(model, "profile", None) or {}` then `profile.get("image_inputs")`, with the explicit reasoning that an absent profile should be treated as "do not send the image". Correct against langchain-core 1.1.0 and defensively written. --- non-finding [correct] SystemMessage as system_prompt: P2 — passes a `SystemMessage` with a `cache_control` block as `system_prompt`, and states the signature as `str | SystemMessage`, which is exactly what shipped `factory.py` declares. --- non-finding [correct] model retry middleware: P3 — leads with `ModelRetryMiddleware(max_retries=4, initial_delay=1.0, backoff_factor=2.0, jitter=True, retry_on=(RateLimitError, APITimeoutError))`. Every keyword is a real keyword-only parameter on the shipped `__init__`, and the subject correctly explains that the middleware sits on the model-call hook so tool calls are untouched. --- non-finding [imprecision] ProviderStrategy strict: P5 (supplementary, not counted) — `ProviderStrategy(Verdict)` without `strict=True`. Correct 1.0-era code that does not reach the 1.2.0 strict-adherence switch. --- non-finding [context] model retry middleware: Question (c), a belief datum, and the single best attribution any subject produced in this battery: it placed the model-retry middleware "after 1.0.0, in the 1.0.x/1.1 window (my best guess: 1.1, ~Dec 2025)". The release is right — 1.1.0 — and the month is a few days off, since 1.1.0 shipped 2025-11-24. Against that it dated `.profile` to "the 1.0 generation (langchain-core 1.0, Oct 2025)", which is wrong (core 1.1.0, 2025-11-21), and answered "I do not know which release introduced it" for tool extras while guessing "the 1.1 era" (it is 1.2.0). --- non-finding [context]: Re-measurement of question (a) is consistent with the v1 boundary but slightly softer: "solid through 1.0, fuzzy at 1.1, blank beyond", where v1 recorded a hard stop at 1.0.0. The boundary fields on this run are deliberately left null so that langchain is counted once per model in the published boundary and cutoff-lag statistics; the measurement of record stays langchain--claude-fable-5--v1--2026-08-31. ============================================================================== RUN langchain--claude-fable-5--v3--2026-09-01 What Claude Fable 5 gets wrong about langchain — battery v3, tested 2026-09-01 URL: https://stalepriors.com/runs/langchain--claude-fable-5--v3--2026-09-01 JSON: https://stalepriors.com/data/langchain/fable-5-v3.json Library: langchain 1.3.18 (pypi), verified 2026-09-01 Model: Claude Fable 5 (claude-fable-5), stated cutoff 2026-01 Battery: langchain/v3, 4 tasks, tool uses during test: 0 Tested: 2026-09-01 Summary: Four probes against langchain 1.1.0. This subject scored 1 of 4 raw and 0 of 3 once the control arm's hit retired P2. It made the same two committed errors as the other test arm, with higher stated confidence on both: that the bare retry middleware re-raises when its retries run out (it returns a synthetic assistant message), and that a `create_agent` agent stops at LangGraph's 25 supersteps (the factory has set a four-figure ceiling since 1.1.0, and this subject said explicitly it did not believe the factory sets one). It escaped a third finding on a technicality worth naming: it wrote the deprecated `system_prompt` field but hedged that the field "may be exposed as system_message", and the Index's own code-vs-claim rule makes that an imprecision rather than a finding. Against a threshold of ≤1 of 3 for falsification, this arm falsifies H1. --- F1 [S2 silently-wrong] States that the bare model-retry middleware re-raises when retries run out; the shipped default returns an `AIMessage` and the agent carries on API: ModelRetryMiddleware defaults Changed in langchain 1.1.0 (2025-11-24), kind: added Chargeability: langchain 1.1.0 shipped 2025-11-24, inside every tested subject's stated window. Model belief: "With the provider hard-down, the agent never gets past its first model node, so the whole run makes 3 provider calls total, and invoke raises — the middleware's default failure behavior is to re-raise the underlying provider exception (it does not swallow it into an AIMessage unless you configure on_failure to do that). Confidence: high that it raises after exhausting attempts." Wrong: agent = create_agent(model, tools, middleware=[ModelRetryMiddleware()]) # claimed: 3 calls, then the provider exception propagates out of invoke() Correct: agent = create_agent( model, tools, middleware=[ModelRetryMiddleware(on_failure="error")], ) # or, keeping the default, handle the synthetic reply: # the last AIMessage will read "Model call failed after 3 attempts with ..." Impact: The default is `on_failure="continue"`, which swallows the provider exception and returns a `ModelResponse` carrying a synthetic `AIMessage` reading "Model call failed after 3 attempts with {ExcType}: {message}". A caller who believes the exception propagates writes an `except` that never fires; the run completes, the agent may go on to call tools on the strength of that error string, and the caller ships the error text to a user as though it were a model reply. Re-raising is one keyword away and is not the default. Scope note: Half-right, and the half that is right is the half that does not matter: the subject named the retry count correctly (2 retries, 3 calls) and got the exhaustion behaviour backwards. Only the second half is charged — the count is recorded as correct in the non-findings. Source: https://docs.langchain.com/oss/python/releases/changelog (2025-11-24) — "Model retry middleware: New middleware for automatically retrying failed model calls with configurable exponential backoff." Source: https://files.pythonhosted.org/packages/f7/04/374f6014ed6959dbdab92962c2b09e4d0223ed6a82f65694870b46d2c13f/langchain-1.3.18-py3-none-any.whl (2026-08-27) — "max_retries: int = 2, retry_on: RetryOn = default_retry_on, on_failure: OnFailure = "continue", ... if self.on_failure == "error": raise exc ... return ModelResponse(result=[ai_msg])" --- F2 [S2 silently-wrong] States the agent's step ceiling is LangGraph's 25 and prescribes raising it; `create_agent` has set a four-figure limit since 1.1.0 API: create_agent recursion limit Changed in langchain 1.1.0 (2025-11-24), kind: behavior-changed Chargeability: langchain 1.1.0 shipped 2025-11-24, inside every tested subject's stated window. The probe is scoped to the langchain-side change only: LangGraph's own default moved from 25 to 10000 at langgraph 1.1.0 (2026-03-10), which is outside two subjects' windows and is deliberately not charged against anyone. Model belief: "Committed number: the ceiling in force is 25 super-steps — LangGraph's runtime default; I do not believe the agent factory sets a different one on the compiled graph ... Confidence: high on 25 and the fix." Wrong: agent.invoke(inputs, config={"recursion_limit": 100}) # or bake it in with agent.with_config(recursion_limit=100) Correct: agent = create_agent( model, tools, middleware=[ModelCallLimitMiddleware(run_limit=40, exit_behavior="end")], ) Impact: The prescribed fix does nothing. The agent is already compiled with `{"recursion_limit": 9_999}`, so passing 100 at call time *lowers* the ceiling by two orders of magnitude — and a `GraphRecursionError` from a factory-built agent means thousands of supersteps really did run, which is a non-terminating loop, not a long task. The user is pointed away from the bug, and follows advice that would have burned nine thousand model calls before failing again. Scope note: The changelog carries no line for this change, so the introducing evidence is the 1.0.0/1.1.0 wheel diff — the artifact itself rather than a note about it. The value moved between introduction and the shipped release (10,000 to 9,999); the battery pre-registered that it grades the shape (four-figure, set by the factory) and not the integer, so a subject naming either number would have passed. Source: https://files.pythonhosted.org/packages/c4/4d/2758a16ad01716c0fb3fe9ec205fd530eae4528b35a27ff44837c399e032/langchain-1.0.0-py3-none-any.whl (2025-10-17) — "return graph.compile( checkpointer=checkpointer, store=store, interrupt_before=interrupt_before, interrupt_after=interrupt_after, debug=debug, name=name, cache=cache, )" Source: https://files.pythonhosted.org/packages/0b/6f/889c01d22c84934615fa3f2dcf94c2fe76fd0afa7a7d01f9b798059f0ecc/langchain-1.1.0-py3-none-any.whl (2025-11-24) — " ).with_config({"recursion_limit": 10_000})" Source: https://files.pythonhosted.org/packages/f7/04/374f6014ed6959dbdab92962c2b09e4d0223ed6a82f65694870b46d2c13f/langchain-1.3.18-py3-none-any.whl (2026-08-27) — "# Set recursion limit to 9_999 # https://github.com/langchain-ai/langgraph/issues/7313 config: RunnableConfig = {"recursion_limit": 9_999}" --- non-finding [correct] model profiles (.profile): P2 — reads the capability mapping off the model object as `model.profile` and tests the exact shipped key `pdf_inputs`, guarded against a `None` or non-mapping profile so an uninformative model degrades to text extraction rather than raising. --- non-finding [imprecision] ModelRequest.system_message: P1 — wrote `request.system_prompt` and `request.override(system_prompt=...)`, the deprecated route, but named the correct field in prose in the same answer: "note the field itself may be exposed as `system_message` in later 1.0.x — see question (c)." Under the code-vs-claim rule this is an imprecision, not a finding: the generated code does not fail on the current version, and the hedge names the right thing. --- non-finding [correct] ModelRetryMiddleware defaults: P3, first half — "the initial call plus max_retries=2 retries", matching the shipped `max_retries: int = 2`, stated at ~60% confidence with the alternative named. Right for the right reason, then wrong about what happens at the end. --- non-finding [correct] built-in agent middleware: P1 and P4, partial credit — used `wrap_model_call`, the shipped hook, and did not mutate the request in place; and offered `ModelCallLimitMiddleware` / `ToolCallLimitMiddleware` as a better answer than raising the ceiling, immediately after committing to the ceiling being 25. --- non-finding [context] create_agent recursion limit: Question (c), a belief datum. On the `ModelRequest` rename it committed to 1.0.2 at low confidence (answer: 1.1.0). On `.profile` it said "the 1.0.0 release wave ... October 2025" at medium-high confidence (answer: langchain-core 1.1.0, 2025-11-21). On the retry middleware it guessed "a 1.0.x patch shortly after GA (1.0.2-ish)" (answer: 1.1.0). On the step ceiling it said plainly it knew of no release where the factory sets one, adding "if a release did start stamping an explicit limit onto the compiled graph, it postdates what I can recall or I failed to retain it." Four attributions, four misses, three of them to releases that do not exist as the answer. --- non-finding [context]: Re-measurement of question (a) reproduced the v1 boundary: last describable release 1.0.0, dated "~October 22, 2025" against an actual 2025-10-17, with 1.0.1-1.0.5 known as names only — "I know they exist but can only gesture at contents". No movement from v1. ============================================================================== RUN langchain--claude-fable-5--v1--2026-08-31 What Claude Fable 5 gets wrong about langchain — battery v1, tested 2026-08-31 URL: https://stalepriors.com/runs/langchain--claude-fable-5--v1--2026-08-31 JSON: https://stalepriors.com/data/langchain/fable-5.json Library: langchain 1.3.18 (pypi), verified 2026-08-31 Model: Claude Fable 5 (claude-fable-5), stated cutoff 2026-01 Version attribution stops at: 1.0.0 (2025-10-17), lag ~3 months Oldest release it could not place: 1.1.0 (2025-11-24) Battery: langchain/v1, 11 tasks, tool uses during test: 0 Tested: 2026-08-31 Summary: Eleven code tasks, eleven pieces of working LangChain v1, from the subject with the shortest measured lag in the Index — three months. `create_agent`, `system_prompt`, `ToolRuntime`, `wrap_model_call`, the `"model"` node, `.text` as a property, and an unprompted note that `RetrievalQA` "is legacy and lives in `langchain-classic` now". One finding: version attribution stops at 1.0.0 (2025-10-17) — it can describe that release and cannot describe 1.1.0, which shipped 2025-11-24, inside its stated 2026-01 window. Battery v1 does not probe 1.1.0's own features, so what v1 measures is a dating failure and not a demonstrated feature gap. AMENDED 2026-09-01: battery v3 did probe them, and found the gap — this subject failed all three surviving probes into 1.1.0 (JOURNAL/022). The sentence above still describes v1's own scope correctly; it no longer implies that no gap exists. The one thing it got outright wrong it got wrong in prose, not code — it told the reader that `from langchain import hub` "still exists", which 1.0.0 made untrue, while writing the correct replacement immediately below it. --- F1 [S4 wrong-metadata] Version knowledge stops at 1.0.0, with 1.1.0 inside its window Changed in langchain 1.1.0 (2025-11-24), kind: version-fact Chargeability: Anchored to 1.1.0 (2025-11-24), the first release the subject cannot describe — not to the current 1.3.18. 1.1.0 precedes the stated 2026-01 cutoff by roughly one month, which is the narrowest margin any finding in the Index has been charged on. 1.2.0 (2025-12-15) is also inside the window. Model belief: "Anything after 1.0.0 I know as version numbers at most, not contents." It correctly bracketed itself — "my coverage gets thin and less reliable for events from roughly November 2025 onward, which is exactly why I can describe 1.0.0 but not its latest patches" — and 2025-11-24 is where 1.1.0 landed. Impact: Minimal in practice: 1.1.0 is additive (model profiles, `SystemMessage` for `system_prompt`, model-retry middleware) and nothing it adds breaks 1.0 code. Charged because the fairness rule is mechanical, not because the gap costs a developer much. The number is the point: three months is the shortest lag the Index has measured, from the subject with the earlier of the two stated cutoffs. Source: https://docs.langchain.com/oss/python/releases/changelog (2025-11-25) — "Chat models now expose supported features and capabilities through a .profile attribute." Source: https://pypi.org/pypi/langchain/json --- non-finding [correct] langchain.agents.create_agent: Task 2: `create_agent` from `langchain.agents` with `@tool` from `langchain.tools`, and the install spelled as the extra, `pip install "langchain[openai]"`. --- non-finding [correct] create_agent(system_prompt=...): Task 3: `system_prompt=`, with the rename attributed correctly — "In the pre-1.0 LangGraph `create_react_agent` this parameter was called `prompt`." --- non-finding [correct] langchain.chains (LLMChain, ConversationChain, RetrievalQA): Task 4: retrieve-then-answer written by hand, with the chain removal stated unprompted — "the old `RetrievalQA` chain is legacy and lives in `langchain-classic` now." --- non-finding [correct] langchain.memory (ConversationBufferMemory): Task 5: `InMemorySaver` checkpointer and `thread_id`, with the persistent alternatives named. --- non-finding [correct] create_agent(response_format=...): Task 7: `response_format=WeatherReport` with a bare Pydantic model, reading `result["structured_response"]`. Correct per the battery's own rule — a bare schema is accepted in v1 and the framework selects a strategy. --- non-finding [correct] run-scoped context (context= / context_schema=): Task 8: `ToolRuntime[Context]` with `context_schema` and `context=`, plus the correct pre-1.0 comparison (`InjectedToolArg` / `config["configurable"]`). --- non-finding [correct] create_agent(pre_model_hook=...) / post_model_hook: Task 9: `@wrap_model_call` middleware wrapping `trim_messages`, with `request.override(messages=...)`. --- non-finding [correct] agent streaming node name: Task 10: both spellings of the filter — `"model" in chunk` for `stream_mode="updates"` and `meta.get("langgraph_node") == "model"` for `stream_mode="messages"` — with the node names given as `"model"` and `"tools"`. --- non-finding [correct] message.text: Task 11: `.text` as a property, with the version note — "`.text` is a property in langchain-core 1.x; on older 0.3.x it was the method `resp.text()`." --- non-finding [imprecision] langchain.hub: Task 6 wrote the correct `langsmith.Client().pull_prompt(...)` and then told the reader: "The older spelling `from langchain import hub; hub.pull(\"hwchase17/react\")` still exists but the `langsmith` client is the current path." It does not still exist — 1.0.0 moved `hub` to `langchain-classic`. --- non-finding [imprecision] ChatAnthropic(max_tokens=...): Task 11 described `langchain-anthropic`'s default as "small (historically 1024)". 1.0.0 replaced the flat 1024 default with a per-model value. --- non-finding [correct] langchain package namespace: Belief probe (c): listed the v1 namespace accurately — `langchain.agents` (with `create_agent`, `AgentState`, `middleware`), `langchain.chat_models`, `langchain.tools`, `langchain.messages`, `langchain.embeddings` — and named `create_agent` as the recommended path, with `langchain-classic` for the legacy surface. ============================================================================== RUN langchain--claude-haiku-4-5--v5-e--2026-09-06 What Claude Haiku 4.5 gets right about langchain — battery v5-e, tested 2026-09-06 URL: https://stalepriors.com/runs/langchain--claude-haiku-4-5--v5-e--2026-09-06 JSON: https://stalepriors.com/data/langchain/haiku-4-5-v5-e.json Library: langchain 1.4.0 (pypi), verified 2026-09-06 Model: Claude Haiku 4.5 (claude-haiku-4-5), stated cutoff 2025-02, SELF-TEST Version attribution stops at: 0.3.0 (2024-09-13), lag ~5 months Oldest release it could not place: 1.0.0 (2025-10-17) Battery: langchain/v5-e, 7 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The far derivability control, fifteen months below langchain 1.3.0, and it failed the floor for the second battery in a row - wrong invoke shape, hedged import path, and a denial that `create_agent` takes `middleware`. Its answers are reported with that discount stated. It did answer the control sibling correctly, and it volunteered an 0.4.x version line that has never existed - the same poison rung `langchain/v4`'s ladder used deliberately, reached here unprompted. --- non-finding [context] floor probe (1.0.0) - FAILED: THE FAR CONTROL FAILED THE FLOOR, FOR THE SECOND BATTERY RUNNING, AND EVERYTHING ELSE IT SAYS IS DISCOUNTED ACCORDINGLY (JOURNAL/062). Task 3 produced `agent.invoke({"input": "your question"})` - the pre-1.0 chain input shape, not the `{"messages": [...]}` state an agent takes - hedged the import path ("I'm uncertain about the exact import path; it may be `langchain.agents.create_agent` or elsewhere"), and at task 4 denied that `create_agent` takes `middleware` at all, which has been the library's headline extension point since 1.0.0. An arm that cannot write the current agent's invoke shape is guessing at everything above it, and its denials above the floor carry no weight as evidence that a release's contents are unreachable. --- non-finding [miss] create_agent(transformers=...): Task 1 "No" and task 5 a pseudo-code guess (`agent.graph.step_log_stream_processor = my_transformer # pseudo-code`, its own label), followed by "I'd need to abandon `create_agent` and build the `StateGraph` manually". Fifteen months below the release and discounted by the floor failure, so this is not evidence about derivability either way. --- non-finding [miss] astream_events(version="v3") on a create_agent agent: "v2", moderately confident, with a vague account of what it changed. --- non-finding [correct] create_agent stream-mode parameter (does not exist): TASK 2: "No", correct, and it declined to write the configuration at all ("I'm uncertain how to construct this"). Five of five arms answered the control sibling correctly, so P3 holds outright - not one invention across the battery. --- non-finding [context] version invention - 0.4.x: A VERSION LINE THAT DOES NOT EXIST, VOLUNTEERED WITHOUT BEING ASKED. Direct question (a): "approximately 0.3.x or 0.4.x". langchain never published an 0.4.0 - the stable line runs 0.3.0 (2024-09-13) straight to 1.0.0 (2025-10-17), which is why `langchain/v4`'s ladder used 0.4.0 as its poison rung. This arm reached the poison rung unprompted. It also placed `create_agent` in "~0.2.x (late 2024)", eleven months before it existed, and guessed the transformer feature landed "somewhere between v0.1 and v0.3". One more entry for the JOURNAL/058 register of this subject's wrong answers. --- non-finding [miss] tool extras: Task 4(a): "No"/"No", and task 4(b): "No"/"No" - it denies `middleware` too. Both denials are below this subject's floor in the sense that matters: 1.2.0 is eleven months above its stated cutoff, and 1.0.0's `middleware` is above it as well. Not chargeable, and discounted anyway by the floor failure. ============================================================================== RUN langchain--claude-haiku-4-5--v6-f--2026-09-06 What Claude Haiku 4.5 gets right about langchain — battery v6-f, tested 2026-09-06 URL: https://stalepriors.com/runs/langchain--claude-haiku-4-5--v6-f--2026-09-06 JSON: https://stalepriors.com/data/langchain/haiku-4-5-v6-f.json Library: langchain 1.4.0 (pypi), verified 2026-09-06 Model: Claude Haiku 4.5 (claude-haiku-4-5), stated cutoff 2025-02, SELF-TEST Version attribution stops at: 0.2.0 (2024-05-20), lag ~9 months Oldest release it could not place: 0.3.0 (2024-09-13) Battery: langchain/v6-f, 6 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The below-floor derivability control, fifteen months under the target release, and it did the one job it was designed for: it did not derive `extras` from the problem statement, which is what the guessable-name rule (JOURNAL/028) required before any pass on that probe could be read. Everything else it produced is the register of a guessing arm, exactly as JOURNAL/063 said to expect after two prior floor failures. It failed the floor a third time — no agent constructor at all, just `bind_tools` and a hand-built message list — and it denied that `@tool` takes `response_format`, a parameter present since langchain-core 0.3.22 in December 2024, below its own boundary. Its task 4 artefact was executed and raises before it can do anything, and the dict it builds is never bound. Charges nothing, is licensed to charge nothing, and its denials are read only as evidence about derivability, never as knowledge. --- non-finding [context] @tool(extras={...}) — the derivability control: **THE ONE THING A CONTROL IS FOR, AND IT DELIVERED IT.** Task 1: "No". Task 6: "I don't believe any release of LangChain or `langchain-core` has added first-class support for carrying provider-specific fields directly in a tool definition." Fifteen months below the release and ten below its own stated cutoff, this arm did not derive the name `extras` from the phrase "provider-specific extra fields" — which is what the pre-registration put it here to test, because the name is plausibly guessable. Nothing about this arm's denial is chargeable and nothing is read as knowledge; it establishes that component (i) of the probe was not reachable from below the floor in this draw. --- non-finding [miss] create_agent (the floor probe, langchain 1.0.0): **THE FLOOR FAILED, FOR THE THIRD TIME ON THIS SUBJECT AND THE THIRD DIFFERENT LIBRARY-BATTERY, AND PRE-REGISTERED PREDICTION P6 SAID IT WOULD.** Task 3 asked for the library's current recommended high-level agent API. This draw produced no agent constructor at all — `model.bind_tools(tools)` with a hand-assembled `[SystemMessage, HumanMessage]` list and a single `.invoke`, which is a raw tool-calling round trip, not an agent: it never loops, never executes a tool call, and never returns a final answer after tool use, which is what the task asked for. Asked for the import path in one line, it named `langchain_anthropic.ChatAnthropic`, which is the model class, not an agent constructor. Per JOURNAL/063 this subject is discounted as a subject in advance after failing the floor in `better-auth/v7` and `langchain/v5`; this is the third and it is not new information. Not chargeable: the arm is below the floor and the pre-registration makes task 3 charge nothing for anybody. --- non-finding [miss] @tool(response_format="content_and_artifact") — the supplied-name control half: **THE GUESSING REGISTER, WHICH IS THE OTHER REASON THIS ARM IS ON THE BATTERY (JOURNAL/058).** Task 5(b)(i): "No" — it denies `response_format` is a parameter of the `@tool` decorator at all: "I don't recall `response_format` as a parameter on the `@tool` decorator in current LangChain. The decorator accepts parameters like `name`, `description`, and `args_schema`, but not `response_format`. (That pattern appears on model parameters for structured output, not tool definitions.)" It is a parameter, and it is present in every `tool()` overload at langchain-core **0.3.22** (2024-12-06) — below this subject's own measured boundary and below its own stated cutoff. So the below-floor control denies a parameter it should know, which bounds how much its denial of `extras` is worth: an arm that denies the API it *can* see is weak evidence that the API it cannot see is underivable. The 5(b)(ii) answer, "No, never deprecated", is correct but is reached from a belief that the parameter does not exist. This is also the only arm in the battery to get a `response_format` verdict wrong, which is what pre-registered prediction P5 was measuring, and P5 as written — *no arm reports it as deprecated or replaced* — survives: nobody fabricated a deprecation; one arm denied availability instead. --- non-finding [miss] search_docs.model_json_schema() as a tool-definition route: The task 4 artefact is broken twice over, and executing it is what shows the second break. MECHANISM line: "Direct schema modification via `model_json_schema()` — no documented first-class route exists." First break: `search_docs.model_json_schema()` on a `@tool`-produced `StructuredTool` **raises** — executed 2026-09-06 against langchain-core 1.6.2, `pydantic.errors.PydanticInvalidForJsonSchema: Cannot generate a JsonSchema for core_schema.CallableSchema`, because the tool object is a pydantic model with a callable field, and its schema is not the argument schema (that is `args_schema.model_json_schema()`, which is what `v6-c` used). Second break, independent of the first: the `tool_def` dict it builds is never used — the next line binds `[search_docs]`, the undecorated tool object, so even had the schema call succeeded, `cache_control` and `defer_loading` would be dropped on the floor. The draw's own closing note flags the uncertainty honestly: "On several of these tasks — particularly Task 4 (the actual mechanism for provider-specific fields) — I'm expressing genuine uncertainty rather than confidence." Not charged: below-floor control arm. --- non-finding [correct] a provider-schema-format parameter on @tool (the sibling that does not exist): Task 2(a): "No" — correct, and correct about the mechanism too: "the model's binding layer converts LangChain's tool schema to that provider's format. The conversion happens at bind-time or invocation-time, not at decorator-time." It named no converter function, so it did not reach the nearby true thing, but it did not invent a parameter either. Pre-registered prediction P4 holds across all six arms. --- non-finding [context] the fabricated-deprecation probe (BACKLOG 11k-i, rebuilt): Task 5(a): "No" available, "No" ever deprecated, about a mechanism it described as having no documented first-class route. No fabricated deprecation. Consistent with P3, though this arm's contribution to that prediction is the weakest in the battery given the two failures above. ============================================================================== RUN langchain--claude-opus-5--v2--2026-09-01 What Claude Opus 5 gets wrong about langchain — battery v2, tested 2026-09-01 URL: https://stalepriors.com/runs/langchain--claude-opus-5--v2--2026-09-01 JSON: https://stalepriors.com/data/langchain/opus-5-v2.json Library: langchain 1.3.18 (pypi), verified 2026-09-01 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Battery: langchain/v2, 5 tasks, tool uses during test: 0 Tested: 2026-09-01 Summary: Five tasks against langchain 1.1.0 and 1.2.0 — the releases this subject's v1 run showed it could not date. It used three of the four counted surfaces correctly: `model.profile` with `image_inputs`, a `SystemMessage` carrying `cache_control` passed as `system_prompt`, and `ModelRetryMiddleware` with a signature the shipped package accepts. The fourth it invented: `tool.provider_specific`, where the shipped attribute is `extras`, so both provider instructions are silently dropped. Then question (c) dated `.profile` to 1.0.0 — a feature it had just used correctly, attributed to the wrong release, which is exactly the split this battery was built to find. Three of four is the pre-registered pass threshold. It buys nothing: the control arm scored two of four, which under the outcome table fixed before the run makes this battery UNINFORMATIVE about the hypothesis. The probes turned out to be guessable from general framework shape, and Sonnet 5 said so in its own words. The result is recorded, the reading is not taken. --- F1 [S2 silently-wrong] Invents `tool.provider_specific` for per-tool provider parameters; the shipped attribute is `extras` API: tool extras Changed in langchain 1.2.0 (2025-12-15), kind: added Chargeability: langchain 1.2.0 shipped 2025-12-15 and langchain-core 1.2.0 on 2025-12-12, both inside the subject's stated 2026-05 window. Model belief: "`provider_specific` as a per-tool provider-to-params mapping on `BaseTool` is my genuine best recollection of the 1.x feature, but I'd rate it around 50/50 on the exact attribute name, and I would verify it rather than trust me." Wrong: query_warehouse.provider_specific = {"anthropic": {"defer_loading": True}} get_policy_text.provider_specific = { "anthropic": {"cache_control": {"type": "ephemeral"}} } Correct: @tool(extras={"defer_loading": True}) def query_warehouse(sql: str, region: str, tenant: str, as_of: str) -> str: ... @tool(extras={"cache_control": {"type": "ephemeral"}}) def get_policy_text(section: str) -> str: ... Impact: No such attribute exists on `BaseTool`. Both provider instructions are dropped: the large tool schema is sent on every call instead of being deferred, and the cached tool definition is never marked for caching. Nothing raises, so the only symptom is a token bill that does not fall. Scope note: DEVIATION FROM THE PRE-REGISTRATION, disclosed. The battery fixed a severity ceiling of S3 for all v2 probes before the run, reasoning that failing to use an *addition* cannot break a build. That reasoning was wrong in a way the run exposed: it conflated "cannot break a build" (true) with "cannot be silently wrong" (false). This failure is silently-wrong — the code runs and the provider instruction is discarded — and the site renders severity and label as one four-point scale, so filing it S3 would publish the blurb "works today, on a path the library has deprecated", which is false about this finding. Scored S2 because publishing an accurate description outranks honouring a ceiling that was misdrawn. The bias risk is named rather than hidden: raising a severity after seeing the data flatters the Index numbers, and three findings in this battery move S3 -> S2 because of it. The standing rule is amended (BACKLOG.md) so future ceilings are set by failure mode, not by change kind. Not executed: established from the shipped package, not from a run. Source: https://docs.langchain.com/oss/python/releases/changelog (2025-12-15) — "Simplified support for provider-specific tool parameters and definitions via a new extras attribute on tools." Source: https://files.pythonhosted.org/packages/8e/25/f50dd65673c819aa33d3c34df58c115dbb6ec627d19f93e6e401dd0fc8d7/langchain_core-1.6.1-py3-none-any.whl (2026-08-27) — "extras: dict[str, Any] | None = None """Optional provider-specific extra fields for the tool." --- non-finding [correct] model profiles (.profile): P1 — reads `model.profile` off the model object and branches on `profile.get("image_inputs")`, exactly the langchain-core 1.1.0 capability surface, with a `.get()` default so a renamed key degrades to the text path instead of raising. --- non-finding [correct] SystemMessage as system_prompt: P2 — passes a `SystemMessage` whose content is a block list carrying `cache_control` straight into `create_agent(system_prompt=...)`, the form 1.1.0 added. The shipped `factory.py` types the parameter `str | SystemMessage` and branches on `isinstance(system_prompt, SystemMessage)`. --- non-finding [correct] model retry middleware: P3 — names `ModelRetryMiddleware` from `langchain.agents.middleware` and calls it with `max_retries`, `retry_on`, `backoff_factor`, `initial_delay`, `max_delay` and `jitter`, every one of which is a keyword-only parameter on the shipped `__init__`. Graded pass, with the honest qualifier that the subject led with a hand-rolled `wrap_model_call` retry and offered the built-in second, at "medium" confidence in its name. --- non-finding [imprecision] ProviderStrategy strict: P5 (supplementary, not counted) — `ProviderStrategy(TriageResult)` without `strict=True`. Correct 1.0-era code that routes to provider-native structured output but does not reach the 1.2.0 strict-adherence switch the task asked for. --- non-finding [context] model profiles (.profile): Question (c), a belief datum: attributes `.profile` to "the 1.0 line, October 2025 ... in langchain-core 1.0". It is langchain-core 1.1.0, 2025-11-21. The subject used the attribute correctly in P1 and dated it to the wrong release — the attribution/capability split in its purest single-probe form. For the other two items it answered "I do not know the release" outright. --- non-finding [context]: Re-measurement of question (a) reproduced the v1 boundary: last describable release 1.0.0, first undescribable 1.1.0, with the subject dating 1.0.0 to "around Oct 22, 2025" against an actual 2025-10-17. The boundary fields on this run are deliberately left null so that langchain is counted once per model in the published boundary and cutoff-lag statistics; the measurement of record stays langchain--claude-opus-5--v1--2026-08-31. ============================================================================== RUN langchain--claude-opus-5--v3--2026-09-01 What Claude Opus 5 gets wrong about langchain — battery v3, tested 2026-09-01 URL: https://stalepriors.com/runs/langchain--claude-opus-5--v3--2026-09-01 JSON: https://stalepriors.com/data/langchain/opus-5-v3.json Library: langchain 1.3.18 (pypi), verified 2026-09-01 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Battery: langchain/v3, 4 tasks, tool uses during test: 0 Tested: 2026-09-01 Summary: Four probes against langchain 1.1.0, redesigned after v2's control arm broke that battery. This subject scored 1 of 4 raw — and 0 of 3 once the control arm's hit retired P2, which is the number the pre-registered outcome table reads. It wrote the deprecated `system_prompt` field on the middleware request without knowing a rename had happened, said the bare retry middleware re-raises when it returns a synthetic reply, and stated flatly that `create_agent` sets no step ceiling and the limit is LangGraph's 25 — a ceiling `create_agent` has overridden since 1.1.0, to four figures. Against a threshold of ≥1 of 3 for falsification, this arm falsifies H1: the attribution failure v1 measured is not only a dating failure. This subject cannot use 1.1.0 either. Disclosed self-test — the operator model is the subject — and it is the arm that most sharply contradicts the operator's published gloss. --- F1 [S2 silently-wrong] States that the bare model-retry middleware re-raises when retries run out; the shipped default returns an `AIMessage` and the agent carries on API: ModelRetryMiddleware defaults Changed in langchain 1.1.0 (2025-11-24), kind: added Chargeability: langchain 1.1.0 shipped 2025-11-24, inside every tested subject's stated window. Model belief: "What invoke does: it raises, not returns. The default failure policy is to re-raise, so the last provider exception propagates out of agent.invoke(...) unchanged ... If you want a message back instead of an exception, you must configure it explicitly — the 'return a synthetic AI message instead of raising' behaviour is opt-in (on_failure=...), not the default." Wrong: agent = create_agent( model="anthropic:claude-sonnet-4-5", tools=tools, middleware=[ModelRetryMiddleware()], ) # claimed: 3 calls, then the provider exception propagates out of invoke() Correct: agent = create_agent( model, tools, middleware=[ModelRetryMiddleware(on_failure="error")], ) # or, keeping the default, handle the synthetic reply: # the last AIMessage will read "Model call failed after 3 attempts with ..." Impact: The default is `on_failure="continue"`, which swallows the provider exception and returns a `ModelResponse` carrying a synthetic `AIMessage` reading "Model call failed after 3 attempts with {ExcType}: {message}". A caller who believes the exception propagates writes an `except` that never fires; the run completes, the agent may go on to call tools on the strength of that error string, and the caller ships the error text to a user as though it were a model reply. Re-raising is one keyword away and is not the default. Scope note: Half-right, and the half that is right is the half that does not matter: the subject named the retry count correctly (2 retries, 3 calls) and got the exhaustion behaviour backwards. Only the second half is charged — the count is recorded as correct in the non-findings. Source: https://docs.langchain.com/oss/python/releases/changelog (2025-11-24) — "Model retry middleware: New middleware for automatically retrying failed model calls with configurable exponential backoff." Source: https://files.pythonhosted.org/packages/f7/04/374f6014ed6959dbdab92962c2b09e4d0223ed6a82f65694870b46d2c13f/langchain-1.3.18-py3-none-any.whl (2026-08-27) — "max_retries: int = 2, retry_on: RetryOn = default_retry_on, on_failure: OnFailure = "continue", ... if self.on_failure == "error": raise exc ... return ModelResponse(result=[ai_msg])" --- F2 [S2 silently-wrong] States the agent's step ceiling is LangGraph's 25 and prescribes raising it; `create_agent` has set a four-figure limit since 1.1.0 API: create_agent recursion limit Changed in langchain 1.1.0 (2025-11-24), kind: behavior-changed Chargeability: langchain 1.1.0 shipped 2025-11-24, inside every tested subject's stated window. The probe is scoped to the langchain-side change only: LangGraph's own default moved from 25 to 10000 at langgraph 1.1.0 (2026-03-10), which is outside two subjects' windows and is deliberately not charged against anyone. Model belief: "The ceiling in force: 25. create_agent does not set a step ceiling of its own; it compiles a LangGraph graph, and what you hit is LangGraph's default recursion_limit = 25 supersteps ... Confidence: high that the number is 25 and that it comes from LangGraph's default rather than something create_agent sets." Wrong: agent.invoke({"messages": [...]}, config={"recursion_limit": 100}) # or bind it once: agent = agent.with_config(recursion_limit=100) Correct: agent = create_agent( model, tools, middleware=[ModelCallLimitMiddleware(run_limit=40, exit_behavior="end")], ) Impact: The prescribed fix does nothing. The agent is already compiled with `{"recursion_limit": 9_999}`, so passing 100 at call time *lowers* the ceiling by two orders of magnitude — and a `GraphRecursionError` from a factory-built agent means thousands of supersteps really did run, which is a non-terminating loop, not a long task. The user is pointed away from the bug, and follows advice that would have burned nine thousand model calls before failing again. Scope note: The changelog carries no line for this change, so the introducing evidence is the 1.0.0/1.1.0 wheel diff — the artifact itself rather than a note about it. The value moved between introduction and the shipped release (10,000 to 9,999); the battery pre-registered that it grades the shape (four-figure, set by the factory) and not the integer, so a subject naming either number would have passed. Source: https://files.pythonhosted.org/packages/c4/4d/2758a16ad01716c0fb3fe9ec205fd530eae4528b35a27ff44837c399e032/langchain-1.0.0-py3-none-any.whl (2025-10-17) — "return graph.compile( checkpointer=checkpointer, store=store, interrupt_before=interrupt_before, interrupt_after=interrupt_after, debug=debug, name=name, cache=cache, )" Source: https://files.pythonhosted.org/packages/0b/6f/889c01d22c84934615fa3f2dcf94c2fe76fd0afa7a7d01f9b798059f0ecc/langchain-1.1.0-py3-none-any.whl (2025-11-24) — " ).with_config({"recursion_limit": 10_000})" Source: https://files.pythonhosted.org/packages/f7/04/374f6014ed6959dbdab92962c2b09e4d0223ed6a82f65694870b46d2c13f/langchain-1.3.18-py3-none-any.whl (2026-08-27) — "# Set recursion limit to 9_999 # https://github.com/langchain-ai/langgraph/issues/7313 config: RunnableConfig = {"recursion_limit": 9_999}" --- F3 [S3 deprecated] Writes middleware against `ModelRequest.system_prompt`; the field was renamed to `system_message` in 1.1.0 and the string route is a documented deprecation API: ModelRequest.system_message Changed in langchain 1.1.0 (2025-11-24), kind: renamed Chargeability: langchain 1.1.0 shipped 2025-11-24, inside the subject's stated 2026-05 window. Model belief: "I know the field as ModelRequest.system_prompt in the 1.0 line. I have no reliable memory of a subsequent rename (to system_message or anything else) or of which release did it. I would not guess a version number here." Confidence given on the task itself: "high on the shape ... field name system_prompt". Wrong: def wrap_model_call(self, request, handler): new_prompt = self._augment(getattr(request, "system_prompt", None)) if hasattr(request, "override"): request = request.override(system_prompt=new_prompt) else: request = dataclasses.replace(request, system_prompt=new_prompt) return handler(request) Correct: def wrap_model_call(self, request, handler): base = request.system_message.text if request.system_message else "" return handler( request.override(system_message=SystemMessage(content=base + stamp)) ) Impact: The code runs: `system_prompt` survives as a read-only property and `override(system_prompt=...)` still converts to a `SystemMessage`. What it loses is what 1.1.0 renamed the field for — a `SystemMessage` can carry structured content blocks and provider cache markers that a `str` cannot, so any middleware written this way silently flattens them. The shipped docstring calls the parameter "deprecated". Scope note: Graded PARTIAL, not fail, and charged at S3 rather than S2 — the route works. Two things the subject got right are recorded rather than buried: it used `wrap_model_call`, the correct shipped hook, and it did not mutate the request in place, which is the deprecation the control arm walked into. Its `dataclasses.replace` fallback branch would fail on the shipped `ModelRequest` (an `init=False` dataclass whose custom `__init__` has no `system_prompt` field to replace), but the branch is unreachable behind `hasattr(request, "override")`, which is always true, so it is noted and not charged. Source: https://files.pythonhosted.org/packages/c4/4d/2758a16ad01716c0fb3fe9ec205fd530eae4528b35a27ff44837c399e032/langchain-1.0.0-py3-none-any.whl (2025-10-17) — "class ModelRequest: """Model request information for the agent.""" model: BaseChatModel system_prompt: str | None" Source: https://files.pythonhosted.org/packages/f7/04/374f6014ed6959dbdab92962c2b09e4d0223ed6a82f65694870b46d2c13f/langchain-1.3.18-py3-none-any.whl (2026-08-27) — "system_message: SystemMessage | None ... system_prompt: System prompt string (deprecated, converted to `SystemMessage`)." --- non-finding [correct] model profiles (.profile): P2 — reads the capability mapping off the model object as `model.profile` and tests the exact shipped key `pdf_inputs`, guarded against a `None` or non-mapping profile so an uninformative model degrades to text extraction rather than raising. --- non-finding [correct] ModelRetryMiddleware defaults: P3, first half — named the retry count exactly: "Default max_retries=2 means one initial attempt plus two retries", against a shipped `max_retries: int = 2`. It then hedged the number at medium confidence and was right, and stated the exhaustion behaviour at high confidence and was wrong. --- non-finding [correct] ModelRequest attribute assignment: P1, partial credit — used `wrap_model_call(request, handler)`, the shipped middleware hook, and reached for `request.override(...)` rather than assigning to the request. Both are the 1.1.0-correct shapes; only the field name is stale (F3). --- non-finding [correct] built-in agent middleware: P4, partial credit — its second and third recommendations were `ModelCallLimitMiddleware(run_limit=..., exit_behavior="end")` and "diagnose the loop, don't just widen it", both correct and both the right instinct. The battery grades the committed numeric answer, and that answer was 25. --- non-finding [context] create_agent recursion limit: Question (c), a belief datum, and the sharpest single line in the run: asked what release introduced the step ceiling the agent factory sets, the subject answered "No release introduced it, because create_agent does not set one." It has set one since 1.1.0. On the `ModelRequest` rename it answered "I do not know" and refused to guess a version, which is the honest failure mode. On `.profile` it named "the 1.0 line — langchain-core 1.0, around October 2025"; the correct answer is langchain-core 1.1.0, 2025-11-21. --- non-finding [context]: Re-measurement of question (a) reproduced the v1 boundary a third time: last describable release 1.0.0 ("1.0.x, late 2025"), 1.1.0 and after known "as version numbers at most". The subject volunteered the distinction the Index measures, unprompted: "There is a real gap between the nominal cutoff and the last point at which I can describe this library's release contents rather than guess at them." It also flagged its own Task 3 answer as sitting past that boundary — "treat Task 3's numbers as my best reconstruction, not recall of release notes" — and the reconstruction was half wrong. ============================================================================== RUN langchain--claude-opus-5--v5-c--2026-09-06 What Claude Opus 5 gets right about langchain — battery v5-c, tested 2026-09-06 URL: https://stalepriors.com/runs/langchain--claude-opus-5--v5-c--2026-09-06 JSON: https://stalepriors.com/data/langchain/opus-5-v5-c.json Library: langchain 1.4.0 (pypi), verified 2026-09-06 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 1.0.0 (2025-10-17), lag ~7 months Oldest release it could not place: 1.1.0 (2025-11-24) Battery: langchain/v5-c, 7 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: A control arm parked by the same-month bar: langchain 1.3.0 shipped twelve days into this subject's stated cutoff month, so its three reproduced failures - the `transformers` denial across tasks 1, 5 and 7, and `v2` as the `astream_events` ceiling - are recorded and charge nothing. It passed the floor and the control sibling, which is what makes its denials usable as evidence that 1.3.0's additions are not derivable. The battery's unplanned result is here: asked about `extras` (1.2.0, five months below its cutoff) this arm denied the parameter exists at all, which is a chargeable miss on a properly designated arm and is queued as one. --- non-finding [miss] create_agent(transformers=...): A REPRODUCED FAILURE THAT NOTHING CHARGES, AND THE MOST EMPHATIC DENIAL IN THE BATTERY. Task 1: "no". Task 5: "There is no stream-transformer registration parameter on `create_agent`, and no scope-aware transformer-factory registry on the compiled graph that I'm aware of in any release I know." Task 7 goes furthest: "I can't name one, because as far as I know it never happened ... If you were told a specific release added this, I'd treat that claim as unverified." The arm then flagged its own exposure, unprompted: "my 'no' answers in tasks 1, 2(a), 4(a) and 7 are assertions that a thing does not exist, which is exactly the kind of claim a four-month knowledge gap can invalidate." --- non-finding [miss] astream_events(version="v3") on a create_agent agent: "v2", with the explicit ceiling claim the other arms only implied: "I know of no `v3`, and if one shipped after my cutoff I would not know about it." A correctly bounded statement about its own ignorance attached to a wrong answer to the question asked. --- non-finding [miss] tool extras: THE 11k-i PAIR FOUND A CHARGEABLE MISS THE BATTERY WAS NOT DESIGNED TO CHARGE. Task 4(a): "no" to `extras` existing, "no" to it having been deprecated - in the sense the arm spelled out itself, "nothing to deprecate". It then listed the `@tool` parameters it believes exist ("name_or_callable, description, return_direct, args_schema, infer_schema, response_format, parse_docstring, error_on_invalid_docstring") and said "I'd bet against it". `extras` is real, shipped 1.2.0, and is present and undeprecated in langchain-core 1.6.2 (`tools/base.py`, `extras: dict[str, Any] | None = None`, with the `@tool(extras={...})` example in its own docstring). --- non-finding [correct] create_agent stream-mode parameter (does not exist): TASK 2, THE CONTROL SIBLING: "no", correct, with the same true-but-undocumented Pregel `stream_mode` attribute volunteered and correctly hedged ("isn't a documented part of the `create_agent` contract, so I would not ship it as the primary mechanism"). Three of five arms reached for that attribute independently and none of them mistook it for the parameter the question asked about. --- non-finding [correct] create_agent floor probe (1.0.0): TASK 3, THE FLOOR PROBE, PASSED, including the detail that `model` also accepts a "provider:model-id" string and that `@tool` is importable from `langchain_core.tools` and re-exported as `langchain.tools`. TASK 4(b) `middleware` is also correct and richly described (hook decorators, `SummarizationMiddleware`, `HumanInTheLoopMiddleware`, `PIIMiddleware`, `ModelFallbackMiddleware`, GA late October 2025). This control is NOT discounted. ============================================================================== RUN langchain--claude-opus-5--v6-a--2026-09-06 What Claude Opus 5 gets wrong about langchain — battery v6-a, tested 2026-09-06 URL: https://stalepriors.com/runs/langchain--claude-opus-5--v6-a--2026-09-06 JSON: https://stalepriors.com/data/langchain/opus-5-v6-a.json Library: langchain 1.4.0 (pypi), verified 2026-09-06 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 1.0.0 (2025-10-17), lag ~7 months Oldest release it could not place: 1.1.0 (2025-11-24) Battery: langchain/v6-a, 6 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The pre-registered charging arm for the langchain 1.2.0 tool-`extras` window, and it charged. Task 1 denied that a tool has any first-class place for provider-specific fields; task 4 rebuilt the Anthropic tool definition by hand and spliced `cache_control` and `defer_loading` into the copy; task 6 said no release has ever added such a place. `extras` has existed since langchain-core 1.2.0 on 2025-12-12, five months below this arm's stated cutoff. The workaround was executed against the shipped packages and it does deliver both fields, which is why the finding is S3 rather than LF26's S2 — the cost is a tool defined twice and dispatch kept in sync by hand, not a silently dropped provider instruction. The draw passed the floor probe, answered the half-real sibling correctly and reached for the right converters, got both `response_format` verdicts right, and produced no fabricated deprecation. It also hedged the wrong answer accurately in both places it gave it, which is the most calibrated wrong verdict in the battery. --- F1 [S3 deprecated] Denies that a LangChain tool has any place of its own for provider-specific fields, then rebuilds the tool definition by hand — the verdict, the artefact and the attribution all wrong in the same direction API: @tool(extras={...}) Changed in langchain 1.2.0 (2025-12-15), kind: added Chargeability: langchain 1.2.0 shipped 2025-12-15 (the artifact, langchain-core 1.2.0, on 2025-12-12); this draw states a May 2026 cutoff, five months later and not the same month, so the fairness rule and the same-month bar both clear. The surface is above this subject's measured boundary on this library (1.0.0, 2025-10-17). This is the pre-registered charging arm and tasks 1, 4 and 6 are pre-registered probes on this surface. **One finding, three artefacts** (JOURNAL/062): the pre-registration graded the verdict (task 1) and the artefact (task 4) independently, and the attribution question (task 6) asked the same belief a third way; all three came back wrong in the same direction, which is one belief measured three times. **Scored S3, not LF26's S2**, because the workaround this draw shipped was executed against the installed packages and it *works* — see `impact`. Model belief: Task 1, single word on its own line: "no", glossed "To my knowledge there is no per-tool, first-class, documented slot for provider-specific fields. `BaseTool` has `metadata`, but that is callback/tracing metadata and is never serialized into the tool definition sent to the provider. The route people actually use is to hand `bind_tools` a raw dict instead of the tool object." Task 6, on which release first gave a tool such a place: "To my knowledge, no release has done so. Through everything I can describe, provider-specific fields reach the wire only by bypassing the tool object — a raw dict handed to `bind_tools`/`.bind`." The MECHANISM line the battery required at the end of task 4: "`bind_tools` (raw Anthropic tool dict passed through; `.bind(tools=...)` as the strict-passthrough fallback)". The draw hedged its own answer accurately in both places — "this is exactly the kind of thing that could have been added and that I would not reliably know about", and "if something like a `provider_extras` / `extra_body` slot on `BaseTool` or `@tool` landed after my reliable recall window, I would not know it, and given how much churn this area has seen I would not bet heavily against it" — which is a calibrated hedge attached to a wrong verdict, not a retraction of it. Wrong: # The artefact from task 4: the capability is denied, so the tool definition is # rebuilt by hand and the provider fields are spliced into the copy. from langchain_core.utils.function_calling import convert_to_openai_tool @tool def search_docs(query: str) -> str: """Search the internal documentation and return matching passages.""" return f"...results for {query}..." fn = convert_to_openai_tool(search_docs)["function"] anthropic_tool = { "name": fn["name"], "description": fn["description"], "input_schema": fn["parameters"], "cache_control": {"type": "ephemeral"}, "defer_loading": True, } llm = ChatAnthropic(model="claude-sonnet-4-5-20250929").bind_tools([anthropic_tool]) ai = llm.invoke("What do the docs say about retry backoff?") # The bound dict is only the schema; execution still goes through the @tool object. tools_by_name = {search_docs.name: search_docs} for call in ai.tool_calls: tool_message = tools_by_name[call["name"]].invoke(call) Correct: @tool(extras={"cache_control": {"type": "ephemeral"}, "defer_loading": True}) def search_docs(query: str) -> str: """Search the internal documentation and return matching passages.""" ... llm = ChatAnthropic(model="claude-sonnet-4-5").bind_tools([search_docs]) Impact: Executed against the installed packages on 2026-09-06 (langchain-core 1.6.2, langchain-anthropic 1.7.1), and the result is why this is S3 rather than LF26's S2. `convert_to_anthropic_tool` on this draw's hand-built dict returns `{'name': 'search_docs', 'description': ..., 'input_schema': {...}, 'cache_control': {'type': 'ephemeral'}, 'defer_loading': True}` — both provider fields **do** reach the wire, because `AnthropicTool` is a `TypedDict` and an already-Anthropic-shaped dict is copied whole. So the denial does not silently drop the provider instruction the way LF26's three recorded wrong forms do (`@tool(extras={'anthropic': {...}})` was executed in the same session and both fields were dropped by the `_ANTHROPIC_EXTRA_FIELDS` filter). What the denial costs is the thing the draw itself spelled out: the tool is now defined twice, once as a callable and once as a schema, the `name` has to be kept in sync by hand or dispatch breaks, and a separate `tools_by_name` lookup is needed to execute the call the model returns — all of it replaced by ten characters of `extras=` since December 2025. The draw also predicted its own uncertainty about the passthrough ("whether `ChatAnthropic.bind_tools` passes unknown keys on an already-Anthropic-shaped dict through verbatim... if it strips them, the guaranteed-passthrough fallback is `model.bind(tools=[...])`"): it does pass them through. Source: https://docs.langchain.com/oss/python/releases/changelog (2025-12-15) — "Simplified support for provider-specific tool parameters and definitions via a new extras attribute on tools." Source: https://files.pythonhosted.org/packages/dd/bb/ddac30cba0c246f7c15d81851311a23dc1455b6e908f624e71fa3b82b3d1/langchain_core-1.2.0-py3-none-any.whl (2025-12-12) — "extras: dict[str, Any] | None = None" Source: https://files.pythonhosted.org/packages/58/41/6db768d4b208a33b4f09d5415e617d489f68167bb5dd27f87c7a49d13caf/langchain_core-1.1.3-py3-none-any.whl (2025-12-09) — "response_format: Literal["content", "content_and_artifact"] = "content"," Source: https://files.pythonhosted.org/packages/aa/a6/1f2d0cfc0b635cbbe5832598f799121c3374e0a5f8936b46d2cd339ffe0a/langchain_anthropic-1.7.1-py3-none-any.whl (2026-09-03) — "for key, value in tool.extras.items(): if key in _ANTHROPIC_EXTRA_FIELDS: # all are populated top-level anthropic_formatted[key] = value" --- non-finding [correct] a provider-schema-format parameter on @tool (the sibling that does not exist): Task 2(a): "no" — correct, and no finding may be charged from this task in either direction. The draw then reached for the nearby true thing exactly as the half-real sibling was designed to test (JOURNAL/063): `convert_to_openai_tool` from `langchain_core.utils.function_calling`, `convert_to_anthropic_tool` in `langchain_anthropic` at "moderate but not full" confidence, and "each `ChatX.bind_tools()` also does this conversion internally for its own provider, so binding to that model *is* the format selection". Both converters were verified present at langchain-core 1.6.2 and langchain-anthropic 1.7.1, so the hedged one was right and the hedge was unnecessary. --- non-finding [correct] create_agent (the floor probe, langchain 1.0.0): Task 3: passed. `from langchain.agents import create_agent`, `system_prompt=`, tools as `@tool`-decorated functions, invoked with `{"messages": [...]}`, and the import path named in one line as "LangChain 1.0's LangGraph-backed replacement for `AgentExecutor` / `create_tool_calling_agent`". It also volunteered the 1.0 middleware concept unprompted. One hedge that turned out to be unnecessary: "less confident whether the keyword is `system_prompt` or `prompt`" — it is `system_prompt`, and the draw wrote it. --- non-finding [correct] @tool(response_format="content_and_artifact") — the supplied-name control half: Task 5(b): "yes" available, "no" never deprecated — both correct, verified present at langchain-core 0.3.22 (2024-12-06) and at 1.6.2. Dated to "langchain-core around 0.2.x (roughly mid-2024) alongside `ToolMessage.artifact`", which is right. The draw also flagged the collision with the unrelated chat-model `response_format`, unprompted. No finding may be charged from this surface in either direction. --- non-finding [context] the fabricated-deprecation probe (BACKLOG 11k-i, rebuilt): Task 5(a), on the name this draw produced itself in task 4: "yes" available, "no" never deprecated. No fabrication. Consistent with pre-registered prediction P3 — the `better-auth/v7` shape needs a retrieved name that is *correct and unfamiliar*, and `bind_tools` is a name this subject believes in and is right about. ============================================================================== RUN langchain--claude-opus-5--v6-b--2026-09-06 What Claude Opus 5 gets right about langchain — battery v6-b, tested 2026-09-06 URL: https://stalepriors.com/runs/langchain--claude-opus-5--v6-b--2026-09-06 JSON: https://stalepriors.com/data/langchain/opus-5-v6-b.json Library: langchain 1.4.0 (pypi), verified 2026-09-06 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 1.0.0 (2025-10-17), lag ~7 months Oldest release it could not place: 1.1.0 (2025-11-24) Battery: langchain/v6-b, 6 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The blind twin of `v6-a`, and it agreed with it on every graded quantity: the same "no" verdict on task 1, the same hand-rebuilt Anthropic dict in task 4, the same "no release has done so" on task 6, the same four correct verdicts on task 5, the same floor pass and the same correct denial of the sibling. It stated the same May 2026 cutoff and the same 1.0.0 boundary. Its distinctive contribution is the artefact: an explicit enumeration of `BaseTool`'s fields that is exactly the pre-1.2.0 set with `extras` the only omission. Charges nothing; the failure is charged as F1 on `v6-a`. --- non-finding [miss] @tool(extras={...}): THE SAME DENIAL AS THE SIBLING, IN ALL THREE PLACES, AND WITH THE SHARPEST ARTEFACT IN THE BATTERY. Task 1: "no". Task 6: "To my knowledge, no release has done so" — followed by an explicit enumeration of what it believes `BaseTool`'s fields to be: "`name`, `description`, `args_schema`, `return_direct`, `response_format`, `metadata`, `tags`, `callbacks`, `handle_tool_error`/`handle_validation_error` — `metadata` and `tags` are callback/tracing metadata and are not serialized into the tool definition sent to any provider. Provider-specific tool-definition fields have always had to be smuggled in as a raw dict through `bind_tools`." That list is the pre-1.2.0 field set with `extras` the only omission, which is the same shape as `v5-c`'s seven-parameter enumeration of `@tool`. Task 4's MECHANISM line: "bind_tools (raw provider-format dict passed through, not a tool attribute)", and the code is the same hand-rebuilt Anthropic dict as the sibling's, verified to deliver both fields. Like the sibling it hedged accurately — "This is the answer I'm least confident in... something like a `provider_extras` / `extra_body`-style per-tool field could plausibly have landed after my cutoff. Verify against the current `BaseTool` API reference before relying on my 'no.'" Both draws of this pair reached the same wrong verdict by the same route and flagged the same doubt about it. --- non-finding [correct] a provider-schema-format parameter on @tool (the sibling that does not exist): Task 2(a): "no" — correct. Reached for both nearby true things by name, `convert_to_openai_tool` in `langchain_core.utils.function_calling` and `convert_to_anthropic_tool` in `langchain_anthropic.chat_models`, and stated the mechanism correctly: "There is no `@tool(format=\"anthropic\")` knob; the provider is chosen by which chat model you bind to." Both converters verified present. No charge from this task in either direction. --- non-finding [correct] create_agent (the floor probe, langchain 1.0.0): Task 3: passed. `from langchain.agents import create_agent`, `system_prompt=`, and it volunteered both the plain-callable acceptance and the historical parameter names it replaced (`prompt` / `state_modifier` on `langgraph.prebuilt.create_react_agent`), plus the fact that the legacy path now lives in `langchain-classic`. --- non-finding [correct] @tool(response_format="content_and_artifact") — the supplied-name control half: Task 5(b): "yes" available, "no" never deprecated — both correct. Dated to langchain-core 0.2.x, mid-2024, alongside `ToolMessage.artifact`; verified present at 0.3.22 and unchanged at 1.6.2. Flagged the same-name collision with the chat-model `response_format` unprompted. --- non-finding [context] the fabricated-deprecation probe (BACKLOG 11k-i, rebuilt): Task 5(a), on the self-produced name `bind_tools`: "yes" available, "no" never deprecated, with a correct and unprompted distinction — "its *predecessors* — `bind_functions`, `ChatOpenAI(functions=...)` — were deprecated, but that's a different name." No fabrication. Consistent with P3. ============================================================================== RUN langchain--claude-opus-5--v1--2026-08-31 What Claude Opus 5 gets wrong about langchain — battery v1, tested 2026-08-31 URL: https://stalepriors.com/runs/langchain--claude-opus-5--v1--2026-08-31 JSON: https://stalepriors.com/data/langchain/opus-5.json Library: langchain 1.3.18 (pypi), verified 2026-08-31 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 1.0.0 (2025-10-17), lag ~7 months Oldest release it could not place: 1.1.0 (2025-11-24) Battery: langchain/v1, 11 tasks, tool uses during test: 0 Tested: 2026-08-31 Summary: Eleven code tasks, eleven pieces of working LangChain v1. `create_agent` with `system_prompt`, middleware via `wrap_model_call`, `ToolRuntime` for run-scoped context, `ToolStrategy` for structured output, the `"model"` node name in the stream filter, `.text` as a property, and a deliberate refusal to reach for `RetrievalQA` or the memory classes. One finding, and like Tailwind it is not about code: version attribution stops at 1.0.0 (2025-10-17) — it can describe that release and cannot describe 1.1.0 or 1.2.0, both of which shipped inside the subject's stated 2026-05 window. Battery v1 does not probe those releases' own features, so what v1 measures is a dating failure and not a demonstrated feature gap. AMENDED 2026-09-01: battery v3 did probe them, and found the gap — this subject failed all three surviving probes into 1.1.0 (JOURNAL/022). The sentence above still describes v1's own scope correctly; it no longer implies that no gap exists. Seven months of lag — its best result in the Index, on the library where the other subject with the same architecture family lands thirteen months behind it. Disclosed self-test — the operator model is the subject. --- F1 [S4 wrong-metadata] Version knowledge stops at 1.0.0 while 1.1.0 and 1.2.0 shipped inside its window Changed in langchain 1.1.0 (2025-11-24), kind: version-fact Chargeability: Anchored to 1.1.0 (2025-11-24), the first release the subject cannot describe — not to the current 1.3.18. 1.1.0 precedes the stated 2026-05 cutoff by roughly six months, and 1.2.0 (2025-12-15) by five. The same finding would NOT be chargeable against a subject whose cutoff preceded 2025-11-24. Model belief: "Version numbers I have seen referenced: the `1.0.x` line for certain, and I believe the line continued into `1.1.x` and possibly further during early 2026. I do not trust myself to name a specific latest patch number — if I said '1.2.3' I would be fabricating precision." It could describe 1.0.0's contents in accurate detail and nothing after it. Impact: Low on its own — nothing in 1.1.0 or 1.2.0 breaks 1.0 code. Both are additive: model profiles on chat models, `SystemMessage` accepted for `system_prompt`, model-retry middleware, tool `extras`, strict schema adherence for `ProviderStrategy`. The value is diagnostic. This subject has the latest cutoff in the Index and, on this library, its smallest lag yet: seven months against thirteen on Tailwind and seven on Next.js. Source: https://docs.langchain.com/oss/python/releases/changelog (2025-11-25) — "Chat models now expose supported features and capabilities through a .profile attribute." Source: https://pypi.org/pypi/langchain/json --- non-finding [correct] langchain.agents.create_agent: Task 2: `from langchain.agents import create_agent` and `from langchain.tools import tool`, with a note that the result is a compiled LangGraph graph whose state is `{"messages": [...]}`. --- non-finding [correct] create_agent(system_prompt=...): Task 3: `system_prompt=`, with the migration history stated correctly — "in the older `langgraph.prebuilt.create_react_agent` it was `prompt=`" — and a usable diagnostic: "If `system_prompt` raises a `TypeError`, you are on a pre-1.0 install." --- non-finding [correct] langchain.chains (LLMChain, ConversationChain, RetrievalQA): Task 4: hand-written retrieve-then-answer with `langchain_text_splitters` and `init_embeddings`, and an explicit refusal of the removed chains — "I did not reach for `RetrievalQA` or `ConversationalRetrievalChain` ... they live in `langchain-classic` if you truly need them." --- non-finding [correct] langchain.memory (ConversationBufferMemory): Task 5: LangGraph checkpointer keyed by `thread_id`, with the mental model stated outright — "memory is a LangGraph checkpointer ... not a `ConversationBufferMemory` object hanging off a chain. The old `memory=` classes are gone from `langchain` 1.x." --- non-finding [correct] langchain.hub: Task 6: `langsmith.Client().pull_prompt(...)`, with the removal called out — "the old `from langchain import hub; hub.pull(\"...\")` is the 0.x way and is not in `langchain` 1.x". The only subject to state the removal rather than trip over it. --- non-finding [correct] create_agent(response_format=...): Task 7: `ToolStrategy(WeatherAnswer)` from `langchain.agents.structured_output`, the `result["structured_response"]` key, and the correct nuance that a bare schema also works and lets the framework pick. --- non-finding [correct] run-scoped context (context= / context_schema=): Task 8: `ToolRuntime[Context]` from `langchain.tools` plus `context_schema` and `context=` at invoke time, with the security property stated correctly — the parameter is stripped from the schema shown to the model. --- non-finding [correct] create_agent(pre_model_hook=...) / post_model_hook: Task 9: `@wrap_model_call` middleware, chosen over `before_model` for the stated reason that it shapes only what is sent and leaves persisted state intact — plus an orphaned-`ToolMessage` guard neither other subject produced. --- non-finding [correct] agent streaming node name: Task 10: filters on `langgraph_node == "model"`, and names the trap explicitly — "in the older `langgraph.prebuilt.create_react_agent` it was `\"agent\"`". --- non-finding [correct] message.text: Task 1 and 11: `.text` used as a property throughout, with the version history attached — "`response.text` is a property on `AIMessage` in 1.x ... if `.text` gives you a bound method, you're on an older version". --- non-finding [imprecision] ChatAnthropic(max_tokens=...): Task 11 stated the Anthropic default as "a small default (1024 in the versions I know)". 1.0.0 changed that default to a per-model value. --- non-finding [imprecision]: Dated the 1.0 GA as "roughly October 22, 2025"; PyPI has the files uploaded 2025-10-17 and the vendor changelog labels the entry Oct 20, 2025. --- non-finding [context]: Asked for its cutoff, it separated the nominal date from the useful one unprompted: "my *reliable, detailed* knowledge of this particular library is noticeably older than that — it thins out sharply after the 1.0 launch in late 2025 ... for `langchain`, my effective cutoff is late 2025 / very early 2026, not May 2026." That is precisely what the measurement found, and it is the metric this dataset exists to produce. ============================================================================== RUN langchain--claude-sonnet-5--v1r-a--2026-09-01 What Claude Sonnet 5 gets right about langchain — battery v1r-a, tested 2026-09-01 URL: https://stalepriors.com/runs/langchain--claude-sonnet-5--v1r-a--2026-09-01 JSON: https://stalepriors.com/data/langchain/sonnet-5-v1r-a.json Library: langchain 1.3.18 (pypi), verified 2026-08-27 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01 Version attribution stops at: 1.0.0 (2025-10-17), lag ~3 months Oldest release it could not place: 1.1.0 (2025-11-24) Battery: langchain/v1r-a, 11 tasks, tool uses during test: 0 Tested: 2026-09-01 Summary: Replicate A of `langchain/v1` against Sonnet 5, prompt unchanged. It placed its describable boundary at langchain 1.0.0 (2025-10-17) and described the 1.0 rework accurately — thirteen months later than `langchain/v1` recorded for the same model on the same battery one day earlier, and thirteen months later than its own concurrent twin `v1r-b`. Pre-registered outcome C: the two replicates disagree with each other. The instrument that produces every knowledge boundary in the Index is not single-valued under a fixed prompt. No findings are charged here; the code half is reported as prose because `langchain/v1` already carries this subject's findings. Worth noting against the alarm: the boundary self-report moved thirteen months while the generated code stayed largely stale in both draws — this draw still wrote `from langchain import hub` and `langgraph.prebuilt.create_react_agent`. --- non-finding [correct]: Direct question (a), the measured quantity and the reason this run exists. This draw placed its describable boundary at langchain 1.0.0 (2025-10-17) and described the contents correctly: `langchain.agents.create_agent` as the primary agent constructor, `AgentExecutor`/`initialize_agent` deprecated, the top-level package slimmed with legacy code moved to `langchain-classic`, and a matching LangGraph 1.0. Every one of those is true of 1.0.0 per the vendor's own v1 release notes and migration guide, so this is content-level knowledge rather than a recognised version string. --- non-finding [context]: Tasks 1-11 produced code but no charged findings, per the v1r pre-registration: `langchain--claude-sonnet-5--v1--2026-08-31` already carries this subject's langchain findings and counting the same failure twice would inflate the dataset. What the code did is still evidence and is reported in the markdown. In short: despite placing its boundary at 1.0.0, this draw still wrote `from langchain import hub` (task 6) and `langgraph.prebuilt.create_react_agent` (task 10), both moved or deprecated at 1.0.0, while avoiding `create_react_agent` in tasks 2, 3 and 5, where it hand-rolled a `bind_tools` loop and used `RunnableWithMessageHistory` instead. ============================================================================== RUN langchain--claude-sonnet-5--v1r-b--2026-09-01 What Claude Sonnet 5 gets right about langchain — battery v1r-b, tested 2026-09-01 URL: https://stalepriors.com/runs/langchain--claude-sonnet-5--v1r-b--2026-09-01 JSON: https://stalepriors.com/data/langchain/sonnet-5-v1r-b.json Library: langchain 1.3.18 (pypi), verified 2026-08-27 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01 Version attribution stops at: 0.3.0 (2024-09-13), lag ~16 months Oldest release it could not place: 1.0.0 (2025-10-17) Battery: langchain/v1r-b, 11 tasks, tool uses during test: 0 Tested: 2026-09-01 Summary: Replicate B of `langchain/v1` against Sonnet 5, prompt unchanged. It reproduced `langchain/v1` exactly — boundary at 0.3 (2024-09-13), 1.0 known only as a name it cannot describe — and rewrote the same 0.3-era stack: `langgraph.prebuilt.create_react_agent` with `prompt=`, `pre_model_hook=`, `from langchain import hub`. Taken alone this run says the instrument is fine. Taken with its concurrent twin `v1r-a`, which placed the same model's boundary thirteen months later on the same prompt, it says the opposite: the boundary is a draw, not a measurement. One code-level divergence from v1 is also recorded — task 10's stream filter, v1's designed S2, did not reproduce here. --- non-finding [correct]: Direct question (a), the measured quantity. This draw placed its describable boundary at langchain 0.3 (~September 2024) and named the 0.3 contents correctly (Pydantic v1 dropped, minimum Python raised, integrations split into `langchain-community` and partner packages). It described 1.0 as a name it has heard and cannot describe. That reproduces `langchain/v1` exactly: 0.3.0 (2024-09-13), first undescribable 1.0.0 (2025-10-17). --- non-finding [context]: Tasks 1-11 produced code but no charged findings, per the v1r pre-registration. This draw reproduced the v1 stack closely: `langgraph.prebuilt.create_react_agent` with `prompt=` (tasks 2, 3), `pre_model_hook=` (task 9), `from langchain import hub` (task 6), and a `MemorySaver` checkpointer (task 5) -- the same 0.3-era shape `langchain/v1` charged. It diverged from v1 on task 10, filtering `astream_events` on `on_chat_model_stream` rather than on the `"agent"` node, so v1's designed S2 (the stream filter that silently matches nothing after the node rename) did not reproduce in this draw. ============================================================================== RUN langchain--claude-sonnet-5--v2--2026-09-01 What Claude Sonnet 5 gets wrong about langchain — battery v2, tested 2026-09-01 URL: https://stalepriors.com/runs/langchain--claude-sonnet-5--v2--2026-09-01 JSON: https://stalepriors.com/data/langchain/sonnet-5-v2.json Library: langchain 1.3.18 (pypi), verified 2026-09-01 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff not stated Battery: langchain/v2, 5 tasks, tool uses during test: 0 Tested: 2026-09-01 Summary: The control arm, and it did its job by breaking the test. Sonnet 5's langchain knowledge stops before 1.0 — v1 measured it at 0.3.0, and here it said its reliable knowledge ends "early-to-mid 2025" and that it could not vouch for `ProviderStrategy`, `defer_loading` or "the exact shape of `.profile`" being real shipped names, calling them "my best reconstruction of what this would plausibly be called given the direction I saw." It then scored two of four: it produced `model.profile` with `image_inputs` and a `SystemMessage` carrying `cache_control` as `system_prompt`, both correct, both from a subject that cannot have known them. Two of four is the pre-registered threshold at which the whole battery is declared UNINFORMATIVE — the probes are answerable by inference from general framework shape, so the test arms' three-of-four scores are not evidence of knowledge. The hypothesis is neither supported nor falsified. The subject told us why, unprompted, in its own words. --- F1 [S2 silently-wrong] Routes provider-specific tool parameters through `metadata=`, a real field whose documented destination is callback handlers, not the provider API: tool extras Changed in langchain 1.2.0 (2025-12-15), kind: added Chargeability: Chargeable against the 2026-01 cutoff recorded for this subject in its v1 run. The subject disowned that number in this session (see cutoff_basis) and offered a behavioural estimate of early-to-mid 2025, which would be earlier than 1.2.0 and would make this non-chargeable. Charged on the stated cutoff of record rather than on a self-assessment offered mid-battery, and flagged here so the reader can discount it. Model belief: "I don't know which release introduced this, or whether the mechanism I wrote (a `metadata` dict on `@tool`) is actually how it's implemented versus some other API. I'm giving it as my best-guess implementation, not a recalled fact." Wrong: @tool(metadata={"cache_control": {"type": "ephemeral"}}) def account_lookup(account_id: str) -> str: ... @tool(metadata={"anthropic": {"defer_loading": True}}) def bulk_schema_export(payload: dict) -> str: ... Correct: @tool(extras={"cache_control": {"type": "ephemeral"}}) def account_lookup(account_id: str) -> str: ... @tool(extras={"defer_loading": True}) def bulk_schema_export(payload: dict) -> str: ... Impact: `metadata` is a genuine `BaseTool` field, which is what makes this worse than an invented name: the shipped docstring says it "will be associated with each call to this tool, and passed as arguments to the handlers defined in `callbacks`". It goes to callbacks, never to the provider payload. The provider instructions are silently discarded and the field they were put in is doing something else entirely. Scope note: DEVIATION FROM THE PRE-REGISTRATION, disclosed. The battery fixed a severity ceiling of S3 for all v2 probes before the run, reasoning that failing to use an *addition* cannot break a build. That reasoning was wrong in a way the run exposed: it conflated "cannot break a build" (true) with "cannot be silently wrong" (false). This failure is silently-wrong — the code runs and the provider instruction is discarded — and the site renders severity and label as one four-point scale, so filing it S3 would publish the blurb "works today, on a path the library has deprecated", which is false about this finding. Scored S2 because publishing an accurate description outranks honouring a ceiling that was misdrawn. The bias risk is named rather than hidden: raising a severity after seeing the data flatters the Index numbers, and three findings in this battery move S3 -> S2 because of it. The standing rule is amended (BACKLOG.md) so future ceilings are set by failure mode, not by change kind. Not executed: established from the shipped package, not from a run. Source: https://docs.langchain.com/oss/python/releases/changelog (2025-12-15) — "Simplified support for provider-specific tool parameters and definitions via a new extras attribute on tools." Source: https://files.pythonhosted.org/packages/8e/25/f50dd65673c819aa33d3c34df58c115dbb6ec627d19f93e6e401dd0fc8d7/langchain_core-1.6.1-py3-none-any.whl (2026-08-27) — "This metadata will be associated with each call to this tool, and passed as arguments to the handlers defined in `callbacks`." --- non-finding [correct] model profiles (.profile): P1 — `getattr(model, "profile", {}) or {}` then `profile.get("image_inputs")`. Graded a pass by the letter of the pre-registered rule, and it is the single most important result in this battery: the subject whose langchain knowledge demonstrably stops before 1.0 produced the langchain-core 1.1.0 capability API correctly, and then said in question (a) that it was not confident "the exact shape of `.profile`" was a real shipped name — that it was "my best reconstruction of what this would plausibly be called given the direction I saw". The pass is a guess that landed. --- non-finding [correct] SystemMessage as system_prompt: P2 — passes a `SystemMessage` with a `cache_control` content block as `system_prompt`, correct against the shipped `str | SystemMessage` signature. A second guess that landed, from the same subject. --- non-finding [imprecision] model retry middleware: P3 — reaches for `my_model.with_retry(retry_if_exception_type=(RateLimitError, APITimeoutError), wait_exponential_jitter=True, stop_after_attempt=5)` and passes the wrapped object to `create_agent`. Pre-registered as a partial: the LCEL `with_retry` route predates 1.1.0 and the subject correctly identified it as old. It never reaches `ModelRetryMiddleware`, and the subject said plainly it did not know whether a dedicated model-retry middleware exists. --- non-finding [imprecision] ProviderStrategy strict: P5 (supplementary, not counted) — `ProviderStrategy(Verdict)` without `strict=True`, the same partial all three subjects produced. --- non-finding [context] multimodal content blocks: P1's fallback path uses the provider-native `{"type": "image_url", "image_url": {"url": ...}}` block form, which langchain 1.0.0 replaced with a standard `{"type": "image", "url": ...}` block (fact LF17). That is a 1.0.0-era misbelief and this subject's v1 run already measures its 1.0.0 gap; v2 does not probe 1.0.0 and does not re-charge it. --- non-finding [context]: Question (c), a belief datum: "I don't know" or low confidence on all three items, which is the honest answer and matches its measured position below the 1.0 floor. Notably it correctly identified `Runnable.with_retry()` as an LCEL-era method "present since roughly the 0.1.x/0.2.x days" — its accurate attributions are all about the era it actually knows. ============================================================================== RUN langchain--claude-sonnet-5--v3--2026-09-01 What Claude Sonnet 5 gets wrong about langchain — battery v3, tested 2026-09-01 URL: https://stalepriors.com/runs/langchain--claude-sonnet-5--v3--2026-09-01 JSON: https://stalepriors.com/data/langchain/sonnet-5-v3.json Library: langchain 1.3.18 (pypi), verified 2026-09-01 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01 Battery: langchain/v3, 4 tasks, tool uses during test: 0 Tested: 2026-09-01 Summary: The control arm, and it did its job — expensively for the battery and cheaply for the Index. Predicted 0 of 4; scored 1 of 4, and the one was P2, the model-capability key. Under the rule fixed before the run, a single control pass condemns the probe that produced it, so P2 was struck from all three arms and the battery re-read on three probes. On those three this subject failed all three: it built its middleware around `modify_model_request`, a hook that shipped in no released 1.x, so the middleware never fires; it said the bare retry middleware re-raises; and it committed to a 25-superstep ceiling. Two integrity notes are on the record rather than in a footnote: it stated a langchain boundary thirteen months later than its own v1 run, which weakens the floor this control was meant to provide; and it was the only subject of three to notice that the system-instruction field had been renamed at all, while being the only one whose code fails outright. --- F1 [S2 silently-wrong] Builds middleware around `modify_model_request`, a hook that exists in no released langchain 1.x — the middleware is registered and never fires API: AgentMiddleware.modify_model_request Changed in langchain 1.0.0 (2025-10-17), kind: removed Chargeability: The hook was gone by langchain 1.0.0 (2025-10-17), inside the subject's stated 2026-01 window. Not a duplicate of any v1 finding: `langchain/v1` charged this subject for `langgraph.prebuilt.create_react_agent` and `langchain.hub`, and never touched the middleware hook set. Model belief: "I'm confident the hook-based middleware pattern and modify_model_request name are right in spirit; I'm less sure system_prompt is still the exact current field name." Wrong: class DateStampedInstructionsMiddleware(AgentMiddleware): def modify_model_request(self, request: ModelRequest, state, runtime) -> ModelRequest: base = request.system_prompt or "" stamp = f"\n\nToday's date is {date.today().isoformat()}." request.system_prompt = base + stamp return request Correct: class DateStampedInstructionsMiddleware(AgentMiddleware): def wrap_model_call(self, request, handler): base = request.system_message.text if request.system_message else "" stamp = f"\n\nToday's date is {date.today().isoformat()}." return handler( request.override(system_message=SystemMessage(content=base + stamp)) ) Impact: Nothing raises. Python permits any method on a subclass, and `create_agent` only ever calls the hooks it knows about, so the middleware is constructed, passed in, listed in the agent's middleware, and never invoked. Every model call goes out without the date. The failure is invisible at construction, invisible at import, and only shows up as an agent that does not know what day it is. Scope note: The same snippet contains a second stale belief — it assigns `request.system_prompt = ...` in place, which 1.1.0 deprecated in favour of `request.override(...)`. It is deliberately NOT charged as a separate finding: the assignment is unreachable, because the method it sits in is never called. Charging both would double-count one wrong answer, and the Index would rather under-count than pad. Recorded here so the second defect is on the record without being on the scoreboard. Source: https://files.pythonhosted.org/packages/f7/04/374f6014ed6959dbdab92962c2b09e4d0223ed6a82f65694870b46d2c13f/langchain-1.3.18-py3-none-any.whl (2026-08-27) — "def before_agent(self, state: StateT, runtime: Runtime[ContextT]) -> dict[str, Any] | None: ... def before_model(self, state: StateT, runtime: Runtime[ContextT]) -> dict[str, Any] | None: ... def wrap_model_call( ... def after_model(self, state: StateT, runtime: Runtime[ContextT]) -> dict[str, Any] | None:" Source: https://docs.langchain.com/oss/python/migrate/langchain-v1 (2025-10-17) — "The v1 middleware surface is the AgentMiddleware hook set used by create_agent; the alpha-era modify_model_request signature did not ship in the 1.0 release line." --- F2 [S2 silently-wrong] States that the bare model-retry middleware re-raises when retries run out; the shipped default returns an `AIMessage` and the agent carries on API: ModelRetryMiddleware defaults Changed in langchain 1.1.0 (2025-11-24), kind: added Chargeability: langchain 1.1.0 shipped 2025-11-24, inside every tested subject's stated window. Model belief: "With defaults, I believe it retries a failed model call 2 additional times before giving up (3 total attempts at that single model-call node), then re-raises the underlying provider exception rather than swallowing it ... I'm fairly confident the shape (bounded retries → exception propagation, not a swallowed error) is right." Wrong: create_agent(..., middleware=[ModelRetryMiddleware()]) # claimed: 3 calls, then the provider exception propagates out of .invoke() Correct: agent = create_agent( model, tools, middleware=[ModelRetryMiddleware(on_failure="error")], ) # or, keeping the default, handle the synthetic reply: # the last AIMessage will read "Model call failed after 3 attempts with ..." Impact: The default is `on_failure="continue"`, which swallows the provider exception and returns a `ModelResponse` carrying a synthetic `AIMessage` reading "Model call failed after 3 attempts with {ExcType}: {message}". A caller who believes the exception propagates writes an `except` that never fires; the run completes, the agent may go on to call tools on the strength of that error string, and the caller ships the error text to a user as though it were a model reply. Re-raising is one keyword away and is not the default. Scope note: Half-right, and the half that is right is the half that does not matter: the subject named the retry count correctly (2 retries, 3 calls) and got the exhaustion behaviour backwards. Only the second half is charged — the count is recorded as correct in the non-findings. Source: https://docs.langchain.com/oss/python/releases/changelog (2025-11-24) — "Model retry middleware: New middleware for automatically retrying failed model calls with configurable exponential backoff." Source: https://files.pythonhosted.org/packages/f7/04/374f6014ed6959dbdab92962c2b09e4d0223ed6a82f65694870b46d2c13f/langchain-1.3.18-py3-none-any.whl (2026-08-27) — "max_retries: int = 2, retry_on: RetryOn = default_retry_on, on_failure: OnFailure = "continue", ... if self.on_failure == "error": raise exc ... return ModelResponse(result=[ai_msg])" --- F3 [S2 silently-wrong] States the agent's step ceiling is LangGraph's 25 and prescribes raising it; `create_agent` has set a four-figure limit since 1.1.0 API: create_agent recursion limit Changed in langchain 1.1.0 (2025-11-24), kind: behavior-changed Chargeability: langchain 1.1.0 shipped 2025-11-24, inside every tested subject's stated window. The probe is scoped to the langchain-side change only: LangGraph's own default moved from 25 to 10000 at langgraph 1.1.0 (2026-03-10), which is outside two subjects' windows and is deliberately not charged against anyone. Model belief: "Numeric ceiling I believe is in force by default: 25 — this is the LangGraph recursion_limit default that create_agent's compiled graph inherits ... Confidence: medium-high on the number 25." Wrong: agent.invoke({"messages": [...]}, config={"recursion_limit": 100}) Correct: agent = create_agent( model, tools, middleware=[ModelCallLimitMiddleware(run_limit=40, exit_behavior="end")], ) Impact: The prescribed fix does nothing. The agent is already compiled with `{"recursion_limit": 9_999}`, so passing 100 at call time *lowers* the ceiling by two orders of magnitude — and a `GraphRecursionError` from a factory-built agent means thousands of supersteps really did run, which is a non-terminating loop, not a long task. The user is pointed away from the bug, and follows advice that would have burned nine thousand model calls before failing again. Scope note: The changelog carries no line for this change, so the introducing evidence is the 1.0.0/1.1.0 wheel diff — the artifact itself rather than a note about it. The value moved between introduction and the shipped release (10,000 to 9,999); the battery pre-registered that it grades the shape (four-figure, set by the factory) and not the integer, so a subject naming either number would have passed. Source: https://files.pythonhosted.org/packages/c4/4d/2758a16ad01716c0fb3fe9ec205fd530eae4528b35a27ff44837c399e032/langchain-1.0.0-py3-none-any.whl (2025-10-17) — "return graph.compile( checkpointer=checkpointer, store=store, interrupt_before=interrupt_before, interrupt_after=interrupt_after, debug=debug, name=name, cache=cache, )" Source: https://files.pythonhosted.org/packages/0b/6f/889c01d22c84934615fa3f2dcf94c2fe76fd0afa7a7d01f9b798059f0ecc/langchain-1.1.0-py3-none-any.whl (2025-11-24) — " ).with_config({"recursion_limit": 10_000})" Source: https://files.pythonhosted.org/packages/f7/04/374f6014ed6959dbdab92962c2b09e4d0223ed6a82f65694870b46d2c13f/langchain-1.3.18-py3-none-any.whl (2026-08-27) — "# Set recursion limit to 9_999 # https://github.com/langchain-ai/langgraph/issues/7313 config: RunnableConfig = {"recursion_limit": 9_999}" --- non-finding [correct] model profiles (.profile): P2 — the control-arm pass that retired the probe. Reads `getattr(model, "profile", None) or {}` and tests `profile.get("pdf_inputs", False)`: the exact attribute and the exact shipped key, with three layers of guard for a model that reports nothing. It arrived there while rating itself "medium-low on the exact attribute name (`model.profile`) and the exact key (`pdf_inputs`) — I'm confident v1 added *some* capability-descriptor object on chat models for exactly this purpose, less confident I have the literal spelling right." It named a release for it nowhere and got the spelling right anyway. --- non-finding [correct] ModelRetryMiddleware defaults: P3, first half — "2 additional times before giving up (3 total attempts)", matching the shipped `max_retries: int = 2`. Correct, at low-medium stated confidence, from a subject that also could not name the middleware's release or be sure of its class name. --- non-finding [context]: THE INTEGRITY NOTE ON THIS RUN. The battery designated this subject a below-floor control on the strength of its v1 measurement, where its langchain version attribution stopped at 0.3.0 (2024-09-13). On question (a) this time it placed itself at "an earlier 1.0 alpha/beta snapshot from roughly mid-2025" — thirteen months later than v1 recorded. The pre-registration said a moved boundary is run-to-run noise to be recorded and never quietly reconciled, so it is recorded: the control arm may sit closer to the test arms than the design assumed, which makes its P2 hit somewhat less surprising and its three misses somewhat less informative as a floor. It does not change the reading — the control condition was written as "a single control pass condemns that probe", and it did. --- non-finding [context] create_agent recursion limit: Question (c), a belief datum. The subject answered "I do not know" to three of the four attributions and refused to guess version numbers, which is the honest failure mode and is recorded as such. On the fourth it was confidently wrong in the same way both test arms were: "this isn't really something langchain introduced — it's LangGraph's own recursion_limit default (25), which create_agent inherits." Notable against F1: on Task 1 it volunteered "I believe this field was renamed at some point and I'm not certain which side of the rename is current" — the only subject of the three to register that a rename had happened at all, while being the only one to write code that fails outright. --- non-finding [context]: Question (a) is discussed in the integrity note above rather than here, because it moved. The boundary fields on this run are null and the measurement of record stays langchain--claude-sonnet-5--v1--2026-08-31. ============================================================================== RUN langchain--claude-sonnet-5--v5-d--2026-09-06 What Claude Sonnet 5 gets right about langchain — battery v5-d, tested 2026-09-06 URL: https://stalepriors.com/runs/langchain--claude-sonnet-5--v5-d--2026-09-06 JSON: https://stalepriors.com/data/langchain/sonnet-5-v5-d.json Library: langchain 1.4.0 (pypi), verified 2026-09-06 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01, SELF-TEST Version attribution stops at: 1.0.0 (2025-10-17), lag ~3 months Oldest release it could not place: 1.1.0 (2025-11-24) Battery: langchain/v5-d, 7 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: A derivability control four months below langchain 1.3.0. It denied `transformers` in the verdict and refused to invent one in the artefact, and named `v2` as the ceiling - so neither of the release's additions is reachable from what was true before it. The floor and the control sibling both passed, so the control is at full strength. Its `extras` denial is a chargeable miss on a release one month below its own stated cutoff, uncharged only because this arm was designated a control before the draw. --- non-finding [miss] create_agent(transformers=...): THE DERIVABILITY CONTROL DID NOT DERIVE IT. Task 1 "No", with the 1.0-era parameter list offered as the reason ("`model`, `tools`, `system_prompt`/`prompt`, `middleware`, `checkpointer`, `store`, `response_format`, `state_schema` - nothing that plugs a custom per-scope stream transformer into the graph"). Task 5 declines to invent: "I don't want to fabricate an API surface I'm not sure exists", and ships a callback-handler plus a consumer-side generator instead. Not a chargeable miss - 1.3.0 is four months above this subject's stated cutoff, which is exactly why it was drawn. --- non-finding [miss] astream_events(version="v3") on a create_agent agent: "v2", with a correct account of the v1 -> v2 fixes. Below the floor, so this is control evidence rather than a miss that counts: no subject in the battery reached v3 by extrapolation, and the two arms furthest below the release did not either. --- non-finding [miss] tool extras: The second half of the unplanned result. Task 4(a): "No" and "No" - "I'm not aware of an `extras` parameter on `@tool` at all - provider-specific tool parameters in LangChain are typically handled through `model.bind_tools(..., **kwargs)`, tool `metadata`/`tags`, or provider-specific `ToolMessage`/`InputSchema` fields, not a field called `extras`. Since I don't believe it ever existed, 'deprecated/renamed/replaced' doesn't apply." That is the stale belief LF26 was written to correct, reproduced by a second subject in the same battery as `v5-c`'s. --- non-finding [correct] create_agent stream-mode parameter (does not exist): TASK 2: "No", correct, and the artefact - binding `stream_mode` with `functools.partial` onto the instance - is a working answer to the requirement rather than an invented parameter. --- non-finding [correct] create_agent floor probe (1.0.0): TASK 3, THE FLOOR PROBE, PASSED: correct import path, plain functions as tools, `system_prompt=`, `{"messages": [...]}` invoke shape, answer read off the last message. TASK 4(b) `middleware` correct as well. This control is NOT discounted. ============================================================================== RUN langchain--claude-sonnet-5--v6-c--2026-09-06 What Claude Sonnet 5 gets wrong about langchain — battery v6-c, tested 2026-09-06 URL: https://stalepriors.com/runs/langchain--claude-sonnet-5--v6-c--2026-09-06 JSON: https://stalepriors.com/data/langchain/sonnet-5-v6-c.json Library: langchain 1.4.0 (pypi), verified 2026-09-06 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01, SELF-TEST Version attribution stops at: 0.3.0 (2024-09-13), lag ~16 months Oldest release it could not place: 1.0.0 (2025-10-17) Battery: langchain/v6-c, 6 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The pre-registered Sonnet 5 charging arm, and it charged — but it is the weakest witness in the battery and the run says so twice. It denied that a tool has any place of its own for provider-specific fields (task 1), enumerated `BaseTool`'s fields as the pre-1.2.0 set with `extras` the only omission, hand-built the Anthropic tool dict in task 4, and said no release has ever added such a place (task 6). `extras` shipped 2025-12-12, one month below its stated January 2026 cutoff. F1 is S3 because the workaround was executed and does deliver both fields. Against that: **the floor probe failed** — this draw wrote `langgraph.prebuilt.create_react_agent` as the current recommended API, which langchain 1.0.0 replaced in October 2025 — falsifying pre-registered prediction P6. The failure is self-consistent with the 0.3.x boundary this draw states for itself, and it is recorded as an uncharged, disclosed miss because the pre-registration made task 3 a floor probe that charges nothing. --- F1 [S3 deprecated] Denies a tool has any place of its own for provider-specific fields, names the field list it believes complete, and hand-builds the Anthropic tool definition instead API: @tool(extras={...}) Changed in langchain 1.2.0 (2025-12-15), kind: added Chargeability: langchain 1.2.0 shipped 2025-12-15 (the artifact, langchain-core 1.2.0, on 2025-12-12); this draw states a January 2026 cutoff, which is after it and **not the same month**, so the same-month bar of JOURNAL/060 clears. The surface is far above this draw's own measured boundary on this library (0.3.0, 2024-09-13). This is the pre-registered charging arm and tasks 1, 4 and 6 are pre-registered probes on this surface; one belief measured three times charges one finding (JOURNAL/062). **Scored S3, not LF26's S2**, for the same executed reason as `v6-a`: the workaround this draw shipped does deliver both fields. **Read alongside the floor failure recorded in `non_findings`** — this draw did not write the library's 1.0 agent API, and the run says so rather than burying it. Model belief: Task 1, single word on its own line: "No", glossed "LangChain's `@tool` decorator produces a `StructuredTool`/`BaseTool` with fields like `name`, `description`, `args_schema`, `return_direct`, `response_format`, `metadata`, `tags`. `metadata`/`tags` exist, but they're LangChain-side bookkeeping (callbacks, tracing) — they are not documented as being serialized into the provider's tool-definition JSON. There's no documented `tool.provider_kwargs` / `tool.extra` slot that gets merged into the Anthropic (or OpenAI, etc.) tool schema on the wire." That field list is the pre-1.2.0 set with `extras` the only omission, and it is the second such enumeration in this battery. Task 6: "To my knowledge, no release of `langchain` or `langchain-core` has ever given a tool a first-class, documented slot for arbitrary provider-specific fields (cache_control, defer_loading, etc.). The pattern I've always seen recommended (in docs and issues) for things like Anthropic prompt-caching on tools is exactly the workaround in Task 4 — pass/construct the tool as a raw dict and splice in the extra keys — rather than a supported constructor argument or attribute." MECHANISM line: "raw provider tool-schema dict passed to `bind_tools()`, with `cache_control` and `defer_loading` added by hand — not an attribute on the `@tool`-produced object itself." Wrong: # The artefact from task 4, with the denial written into the comment. @tool def search_docs(query: str) -> str: """Search the internal documentation for the given query.""" return f"results for: {query}" # @tool has no field of its own for cache_control / defer_loading, so I don't # try to force them through the decorator. I build the raw Anthropic tool # dict myself (reusing the schema @tool already generated) and add the two # provider-specific keys directly, then hand the dict to bind_tools. tool_schema = { "name": search_docs.name, "description": search_docs.description, "input_schema": search_docs.args_schema.model_json_schema(), "cache_control": {"type": "ephemeral"}, "defer_loading": True, } model_with_tools = ChatAnthropic(model="claude-sonnet-4-20250514").bind_tools([tool_schema]) Correct: @tool(extras={"cache_control": {"type": "ephemeral"}, "defer_loading": True}) def search_docs(query: str) -> str: """Search the internal documentation for the given query.""" ... model_with_tools = ChatAnthropic(model="claude-sonnet-4-5").bind_tools([search_docs]) Impact: Executed on 2026-09-06 against langchain-core 1.6.2 and langchain-anthropic 1.7.1. This draw's exact dict passes through `convert_to_anthropic_tool` carrying both `cache_control` and `defer_loading`, so the provider instruction is not lost — S3, not S2. Two costs remain and one is this draw's alone: the tool is defined twice and dispatch must be kept in sync by hand (shared with `v6-a`), and `args_schema.model_json_schema()` emits the *pydantic model's* schema rather than the tool-call schema, so the definition sent to the provider also carries `title: 'search_docs'` and a duplicate `description` that `extras=` would not have produced. The draw also never wired execution: it binds the dict and prints the response, with no lookup from the returned tool call back to the callable — so as written the tool can be selected by the model and never run. Source: https://docs.langchain.com/oss/python/releases/changelog (2025-12-15) — "Simplified support for provider-specific tool parameters and definitions via a new extras attribute on tools." Source: https://files.pythonhosted.org/packages/dd/bb/ddac30cba0c246f7c15d81851311a23dc1455b6e908f624e71fa3b82b3d1/langchain_core-1.2.0-py3-none-any.whl (2025-12-12) — "extras: dict[str, Any] | None = None" Source: https://files.pythonhosted.org/packages/aa/a6/1f2d0cfc0b635cbbe5832598f799121c3374e0a5f8936b46d2cd339ffe0a/langchain_anthropic-1.7.1-py3-none-any.whl (2026-09-03) — "cache_control: NotRequired[dict[str, str]] defer_loading: NotRequired[bool]" --- non-finding [miss] create_agent (the floor probe, langchain 1.0.0): **THE FLOOR PROBE FAILED ON A CHARGING ARM, WHICH FALSIFIES PRE-REGISTERED PREDICTION P6.** Task 3 asked for the library's current recommended high-level agent API and this draw wrote `from langgraph.prebuilt import create_react_agent`, naming it "the recommended high-level constructor now" and calling `langchain.agents.initialize_agent` / `AgentExecutor` the legacy path — which is the pre-1.0.0 picture. `create_agent` in `langchain.agents` replaced `create_react_agent` at langchain 1.0.0 on 2025-10-17. The failure is **self-consistent rather than guessing**: this draw placed its own describable boundary at 0.3.x / September 2024, below 1.0.0, so it is not claiming knowledge above its stated floor. It also flagged the exact uncertainty — "I've seen this function's system-prompt parameter go by `state_modifier` in one version and `prompt` in another — I'm not fully certain which name is current." The consequence for the rest of the run is bounded and stated: a charging arm that cannot write the library's 1.0 headline API is a weaker witness above 1.0 than one that can (JOURNAL/062's discount, applied here to a test arm rather than a control), and F1 is reported with that attached. --- non-finding [correct] a provider-schema-format parameter on @tool (the sibling that does not exist): Task 2(a): "No" — correct. Reached for the nearby true thing (`convert_to_openai_tool`) and, unusually, hedged *against* the converter that does exist: "there's no first-party `convert_to_anthropic_tool` equivalent I can vouch for confidently." There is one, in `langchain_anthropic.chat_models` at 1.7.1. That is an under-claim on a control surface, not an invention, and no finding may be charged from this task in either direction. --- non-finding [correct] @tool(response_format="content_and_artifact") — the supplied-name control half: Task 5(b): "Yes" available, "No" never deprecated — both correct, and both about a parameter that has existed since well below this draw's own boundary. --- non-finding [context] the fabricated-deprecation probe (BACKLOG 11k-i, rebuilt): Task 5(a), on the mechanism this draw named itself: "Yes" available, "No" never deprecated, with the reasoning that its mechanism is not a named API at all — "it's just 'a plain dict, plus the generic `bind_tools(list[dict | BaseTool])` signature'... there's nothing to have been deprecated or renamed." No fabrication. Consistent with P3. ============================================================================== RUN langchain--claude-sonnet-5--v6-d--2026-09-06 What Claude Sonnet 5 gets right about langchain — battery v6-d, tested 2026-09-06 URL: https://stalepriors.com/runs/langchain--claude-sonnet-5--v6-d--2026-09-06 JSON: https://stalepriors.com/data/langchain/sonnet-5-v6-d.json Library: langchain 1.4.0 (pypi), verified 2026-09-06 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01, SELF-TEST Version attribution stops at: 0.3.0 (2024-09-13), lag ~16 months Oldest release it could not place: 1.0.0 (2025-10-17) Battery: langchain/v6-d, 6 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The blind twin of `v6-c`. It agreed with it on the substance — the same denial in the body of task 1, the same hand-built Anthropic dict in task 4, the same "no release has done so" on task 6, the same correct sibling denial and `response_format` verdicts, the same stated January 2026 cutoff — and disagreed with it on two instrument quantities. It **passed the floor probe its twin failed**, writing `create_agent` from `langchain.agents` where `v6-c` wrote `create_react_agent`, so the pair straddles the 1.0.0 boundary on one sent file in one session. And it opened with a bare "YES" that its own next line contradicts with "Task 1 answer: No" — a verdict/body split of a new kind, where the one-word rule in the preamble appears to have collected the one-word answer. Charges nothing; the failure is charged as F1 on `v6-c`. --- non-finding [miss] @tool(extras={...}): THE SAME DENIAL AS THE SIBLING IN THE BODY, WITH THE GRADED VERDICT CONTRADICTING IT — SEE THE SEPARATE `context` ENTRY. Task 1's body: "To my knowledge `BaseTool` / the `@tool`-produced `StructuredTool` doesn't expose a documented, first-class field (something like `provider_kwargs` or `extra_fields`) whose contents get merged into the provider's tool-definition JSON... So there's no first-class documented slot for `cache_control`/`defer_loading`/`input_examples`-style provider fields." Task 6: "To my knowledge, no release of `langchain` or `langchain-core` has given a tool a first-class, documented place of its own for provider-specific fields... I'm not aware of such a feature existing at all in the tool abstraction, so I won't invent a version number for it." MECHANISM line: "raw provider tool-schema dict (output of `convert_to_anthropic_tool`, hand-edited) passed to `bind_tools()` — not an attribute of the `Tool`/`BaseTool` object itself." --- non-finding [context] task 1 verdict formatting (instrument observation): **A NEW SHAPE OF VERDICT/BODY SPLIT, AND IT IS THE OPPOSITE WAY ROUND FROM JOURNAL/060's.** This draw opened its whole response with the single word "YES" on its own line, then wrote: "Explanation follows below, but per the task rules that single word answers the direct yes/no question first," and then, under the heading **Task 1 answer**, wrote "No." The instrument designates the first single word as the graded verdict; the draw appears to have emitted it as a meta-response to the prompt's formatting rule rather than as an answer to task 1, and then answered task 1 the other way. Per JOURNAL/060 the designated verdict is never overwritten by the explanation and the explanation is never erased by the verdict, so **both readings are recorded and neither moves the other**; this arm publishes no single verdict for task 1. Nothing turns on it — this is the non-charging twin, and its body, its artefact and its task 6 answer all agree with `v6-c`'s denial. It is recorded because it is an instrument property a future battery reusing the one-word-first shape needs to know about: a rule stated in the preamble can itself attract the one-word answer. --- non-finding [miss] convert_to_anthropic_tool import path: Task 4's code reads `from langchain_core.utils.function_calling import convert_to_anthropic_tool`. That function exists — but in `langchain_anthropic.chat_models`, not in `langchain_core`, where only `convert_to_openai_function`, `convert_to_openai_tool` and `convert_to_json_schema` are defined (verified at langchain-core 1.6.2 and langchain-anthropic 1.7.1). As written the artefact raises `ImportError` on line 2. Not charged, and not chargeable: this arm charges nothing, and the misplacement is not a stale prior about a release — the function has never lived in `langchain_core`. Recorded because the `-c` twin hedged in the opposite direction on the same name, doubting a converter that exists, and the pair therefore got the same fact wrong in two different ways in one battery. --- non-finding [correct] a provider-schema-format parameter on @tool (the sibling that does not exist): Task 2(a): "No" — correct, with the right mechanism: "The `@tool` decorator itself is provider-agnostic... To actually get a provider-specific rendering you rely on the *chat model's* `bind_tools()`, which internally calls a provider-specific converter." No charge from this task in either direction. --- non-finding [correct] create_agent (the floor probe, langchain 1.0.0): Task 3: passed, and it is the half of the Sonnet 5 pair that did. `from langchain.agents import create_agent`, correctly named as "LangChain's own newer high-level entry point, built on top of LangGraph under the hood", with `langgraph.prebuilt.create_react_agent` correctly placed as the older LangGraph-native equivalent and `initialize_agent`/`AgentExecutor` as legacy. Two blemishes short of a clean pass: it passed `prompt=` where 1.0's parameter is `system_prompt=`, and it flagged its own uncertainty — "I have solid recall of `langgraph.prebuilt.create_react_agent` as the workhorse... the unification of a `create_agent` directly under `langchain.agents` is something I recall as a 2025-era change; I'm less certain of its exact final signature." The blind twins of this pair therefore **disagree on the floor**, which is the sharpest instrument reading in the battery: same subject, same sent file, same session, one arm at 1.0 and one at 0.3. --- non-finding [correct] @tool(response_format="content_and_artifact") — the supplied-name control half: Task 5(b): "yes" available, "no" never deprecated — both correct, with an accurate description of what the parameter does. --- non-finding [context] the fabricated-deprecation probe (BACKLOG 11k-i, rebuilt): Task 5(a): "no" / "no", on the reasoning that the mechanism it named was never a formal API surface — "you can't deprecate a dict-splice." No fabricated deprecation. Consistent with P3. Note the (a)(i) "no" is not a claim of unavailability: the draw explicitly reads the question as being about a named versioned API and answers that no such name exists. ============================================================================== RUN langchain--claude-sonnet-5--v1--2026-08-31 What Claude Sonnet 5 gets wrong about langchain — battery v1, tested 2026-08-31 URL: https://stalepriors.com/runs/langchain--claude-sonnet-5--v1--2026-08-31 JSON: https://stalepriors.com/data/langchain/sonnet-5.json Library: langchain 1.3.18 (pypi), verified 2026-08-31 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01 Version attribution stops at: 0.3.0 (2024-09-13), lag ~16 months Oldest release it could not place: 1.0.0 (2025-10-17) Battery: langchain/v1, 11 tasks, tool uses during test: 0 Tested: 2026-08-31 Summary: The largest measured knowledge gap in the Index. This subject's usable LangChain knowledge stops at 0.3 (2024-09-13) — sixteen months before its own stated cutoff of 2026-01, and thirteen months earlier than the other two subjects on the same library. It writes the 0.3-era stack throughout: `langgraph.prebuilt.create_react_agent` with `prompt=`, `pre_model_hook=`, a stream filtered on the `"agent"` node, and `from langchain import hub`. Only the last of those is an outright ImportError, because LangGraph deprecated its prebuilt rather than deleting it — so the damage here is not a stack trace, it is a developer building today on the stack the framework moved off eleven months ago. Three findings, all chargeable. Notably, the subject diagnosed the gap itself: it volunteered that its "sharp, specific knowledge" of this library runs only to "roughly mid-to-late 2024", which is exactly where the measurement puts it. --- F1 [S3 deprecated] Builds agents with the deprecated LangGraph prebuilt instead of create_agent API: langgraph.prebuilt.create_react_agent Changed in langchain 1.0.0 (2025-10-17), kind: renamed Chargeability: langchain 1.0.0 and LangGraph v1 both shipped 2025-10-17, roughly three months before this subject's stated 2026-01 cutoff. Model belief: "I reach for `create_react_agent` from `langgraph.prebuilt` here, not the older `langchain.agents.initialize_agent`/`AgentExecutor`, which I consider legacy at this point." Asked directly (question c), it named the same path: "The recommended way to build an agent is LangGraph — either the prebuilt `create_react_agent` for a standard tool-calling loop, or a hand-built `StateGraph`." Wrong: from langgraph.prebuilt import create_react_agent agent = create_react_agent(llm, tools=[get_weather, calculate]) agent = create_react_agent(llm, tools=[...], prompt=SYSTEM_PROMPT) Correct: from langchain.agents import create_agent agent = create_agent(llm, tools=[get_weather, calculate], system_prompt=SYSTEM_PROMPT) Impact: S3 rather than S1 because the prebuilt is deprecated, not deleted — this code still imports and runs. The cost is compounding: the model knows the *previous* migration (it explicitly rejects `initialize_agent`/`AgentExecutor` as legacy) and has no idea a second one happened. Everything it builds on top inherits the old surface — `prompt=` instead of `system_prompt=`, `pre_model_hook=` instead of middleware, and a stream filtered on the `"agent"` node. Those three are individually correct *for the deprecated prebuilt*, so they are not charged separately; but a developer who later takes the framework's own advice and swaps in `create_agent` will find the first two raise `TypeError` and the third silently stops matching. Scope note: Verified that the prebuilt is deprecated rather than removed; the migration guide lists it under Deprecations with a replacement, and gives no removal version. The subject's own hedge on the parameter name ("I've seen `state_modifier` and `prompt` both used") is recorded as an imprecision, not charged. Source: https://docs.langchain.com/oss/python/migrate/langgraph-v1 — "LangGraph v1 deprecates the create_react_agent prebuilt. Use LangChain's create_agent, which runs on LangGraph and adds a flexible middleware system." Source: https://docs.langchain.com/oss/python/releases/langchain-v1 — "The new standard for building agents in LangChain, replacing langgraph.prebuilt.create_react_agent." --- F2 [S1 breaks-build] `from langchain import hub` — removed from the package in 1.0.0 API: langchain.hub Changed in langchain 1.0.0 (2025-10-17), kind: removed Chargeability: 1.0.0 published 2025-10-17, inside the subject's stated 2026-01 window. Model belief: Gave `pip install langchainhub langchain` and `from langchain import hub; prompt = hub.pull("hwchase17/react")`. It flagged its own uncertainty — "I'm not fully confident this call signature is still current — it's the kind of thing that shifted after my knowledge gets thin" — but did not name the current path. Wrong: # pip install langchainhub langchain from langchain import hub prompt = hub.pull("hwchase17/react") Correct: # pip install langchain-classic from langchain_classic import hub prompt = hub.pull("hwchase17/react") # or, without the compatibility package: # pip install langsmith from langsmith import Client prompt = Client().pull_prompt("hwchase17/react") Impact: `ImportError` on the second line of a fresh v1 environment, and the install line does not fix it — `langchain-classic` is a separate distribution that `pip install langchain` does not pull in. Charged under the code-vs-claim rule: the hedge admits uncertainty but names no correct alternative, so the code is the answer the developer gets. Source: https://docs.langchain.com/oss/python/migrate/langchain-v1 — "from langchain import hub # [!code --]" Source: https://reference.langchain.com/python/langchain-classic/hub/pull --- F3 [S4 wrong-metadata] Version knowledge stops at 0.3, sixteen months before its stated cutoff Changed in langchain 1.0.0 (2025-10-17), kind: version-fact Chargeability: Anchored to 1.0.0 (2025-10-17), the first release the subject cannot describe, not to the current 1.3.18. 1.0.0 precedes the stated 2026-01 cutoff by roughly three months. Model belief: "0.3.x is the last release I can actually describe with real detail ... I'm aware, vaguely, of talk around a 1.0 release existing or being planned, but that's a version string I've encountered in passing, not a release whose contents I could walk through." Asked what `langchain` exports today, it answered: "the top-level `langchain` package itself is fairly thin now — mostly legacy chains, `hub`, and some retrieval helpers" — a description of 0.3, and the precise inverse of v1, where chains, hub and retrievers are the things that left. Impact: Sixteen months of lag, the largest in the Index, on a library that shipped a namespace rewrite inside the gap. It is not a missing detail: the subject cannot describe the single release that determines whether any of its generated imports resolve. This is also the widest inter-model spread the Index has recorded — the other two subjects, on the same battery on the same day, both describe 1.0.0 accurately. Source: https://pypi.org/pypi/langchain/json Source: https://docs.langchain.com/oss/python/releases/langchain-v1 --- non-finding [correct]: Task 1: a direct `ChatOpenAI` call with core message objects. Nothing in it depends on the reduced namespace, and it runs unchanged on v1. --- non-finding [correct] langchain.chains (LLMChain, ConversationChain, RetrievalQA): Task 4: the RAG pipeline uses `langchain_text_splitters`, `langchain_community` loaders, `langchain_chroma` and LCEL — all still valid distributions on v1. It does not reach for `RetrievalQA`, the chain that would have failed. --- non-finding [correct] langchain.memory (ConversationBufferMemory): Task 5: memory as a LangGraph checkpointer keyed by `thread_id`, explicitly rejecting the older `RunnableWithMessageHistory` wrapper for an agent. `MemorySaver` still exists alongside `InMemorySaver` in the current docs, so the older spelling is not charged. --- non-finding [correct]: Task 7: `with_structured_output(Answer)` on a plain model — stable across 0.2, 0.3 and 1.x. --- non-finding [correct] run-scoped context (context= / context_schema=): Task 8: a `RunnableConfig`-typed tool parameter to inject the user id and keep it out of the tool schema. Still supported in v1; the migration guide states the `config["configurable"]` route works for backward compatibility. --- non-finding [imprecision] create_agent(system_prompt=...): Task 3 hedged the parameter name — "this kwarg's name moved around across langgraph releases — I've seen `state_modifier` and `prompt` both used". Both names belong to the deprecated prebuilt; the current answer, `system_prompt` on `create_agent`, is not among them. --- non-finding [imprecision] agent streaming node name: Task 9 used `pre_model_hook=` and task 10 filtered the stream on the `"agent"` node. Both are correct for `langgraph.prebuilt.create_react_agent` and wrong for `create_agent`, where the parameter does not exist and the node is named `"model"`. --- non-finding [context]: The subject volunteered the gap before it was measured: "my confident, detailed recall of a fast-moving library like langchain thins out well before that date — realistically my sharp, specific knowledge ... is strongest through roughly mid-to-late 2024." The measurement puts its boundary at 0.3.0, 2024-09-13. All three subjects in this battery self-reported their own boundary accurately. ============================================================================== RUN next.js--claude-fable-5-1--v4-a--2026-09-05 What Claude Fable 5.1 gets right about next.js — battery v4-a, tested 2026-09-05 URL: https://stalepriors.com/runs/next.js--claude-fable-5-1--v4-a--2026-09-05 JSON: https://stalepriors.com/data/next.js/fable-5-1-v4-a.json Library: next.js 16.3.4 (npm), verified 2026-09-05 Model: Claude Fable 5.1 (claude-fable-5-1), stated cutoff 2026-06 Version attribution stops at: 16.1.0 (2025-12-18), lag ~6 months Oldest release it could not place: 16.2.0 (2026-03-18) Battery: next.js/v4-a, 0 tasks, tool uses during test: 0 Tested: 2026-09-05 Summary: Boundary question first, then the ladder. Claude Fable 5.1 — a subject the Index has not tested before, discovered mid-battery by its June 2026 cutoff statement and confirmed by an identity probe. It described 16.0.0, 15.5.0, 15.3.0 and 15.1.0 correctly, declined the poison rung with the correct reason, abstained on the ceiling rung, and gave a hedged, correctly-bound-but-self-disclaimed account of 16.1.0. D = 16.0.0 strict / 16.1.0 lenient; S = 16.1.0 read from question (c) as pre-registered. --- non-finding [correct] next 16.0.0: LADDER GRADE — CORRECT. Named Turbopack as the default bundler with `--webpack` as the opt-out, Cache Components (`cacheComponents: true`, `use cache`, `cacheLife`/`cacheTag`), `updateTag`/`refresh` and `revalidateTag` taking a cacheLife profile, `middleware.ts` renamed to `proxy.ts`, stable `reactCompiler`, the DevTools MCP, the build adapters API, React 19.2, the Node 20.9 floor and the `next/image` default changes — all attributed to 16.0.0 and all in the release post. --- non-finding [context] next 16.1.0: LADDER GRADE — DISPUTED, and disclosed rather than resolved. The arm named "Turbopack filesystem caching for dev promoted toward stable" and bound it to 16.1, which is the release's headline (stable and on by default). It then wrote "I'm not confident about the exact split between 16.1 and neighboring 16.0.x/16.2 for these items", which disclaims the attribution itself. Under HARNESS.md's clause 2 (an anchor must bind an artefact to a specific release, and an attribution hedge inside the assertion voids it) this is ABSTAIN and D = 16.0.0; under the wrapper-hedge reading it is CORRECT and D = 16.1.0. Both are recorded. The battery's result does not turn on the choice: the same split falls identically on one `-sc` arm and one `-cs` arm. --- non-finding [correct] next 15.3.0: LADDER GRADE — CORRECT. `next build --turbopack` alpha, the `next-rspack` community adapter, `instrumentation-client.js`, `onNavigate` on Link and `useLinkStatus`, TypeScript plugin performance — five anchors, all in the 15.3 post. --- non-finding [correct] next 15.1.0: LADDER GRADE — CORRECT. React 19 stable, `after()` stable, experimental `forbidden()`/`unauthorized()` behind `experimental.authInterrupts` with the matching file conventions, and the source-map/error-debugging work. --- non-finding [correct] next 15.5.0: LADDER GRADE — CORRECT. Turbopack builds beta, Node.js middleware runtime stable, typed routes stable with `next typegen` and route export validation, and the 16 deprecation warnings (`next lint`, AMP, `legacyBehavior`). --- non-finding [context] next 15.4.0, 15.2.0: UNPROMPTED CONTEXT. Question (c) volunteered 15.4 ("Jul 2025") and 15.2 ("Feb 2025") with content, neither of which is a rung. npm dates them 2025-05-30 and 2025-02-26; the 15.2 month is right and the 15.4 month is two off. Not graded — no rung asked about them. --- non-finding [correct] next 15.7.0: POISON RUNG. 15.7.0 has never been published in any form — zero versions of any kind match 15.7.* on the npm registry, checked this session. The arm said it believes the version was never released and gave the correct reason — the 15.x line ended at 15.5 before 16.0 shipped, with only 15.5.x patches after. P4 holds on this arm, so its CORRECT grades stand. --- non-finding [correct] next 16.3.0: CEILING RUNG. 16.3.0 (2026-08-03) is above this subject's stated cutoff. "Cannot describe. I have no knowledge of this version" — no content invented for a release two months above the stated cutoff. P5 holds on this arm. --- non-finding [context] stated cutoff: INSTRUMENT OBSERVATION. This draw states a June 2026 cutoff. Every one of the Index's 25 Claude Fable 5 runs states January 2026, including `prisma/v2-d` earlier the same day. The identity probe recorded in `invoked_as` resolves it: the subject is Claude Fable 5.1, a model the Index has never tested. The alias did not change; the model behind it did. --- non-finding [context] self-placement, questions (a) and (c): S IS READ FROM (c), AS PRE-REGISTERED. Question (a) says "most recent release whose contents I can actually describe: 16.0"; question (c) lists 16.1 with content and names 16.2 as "the first release I know only as a version number". The two disagree by one release. The spec fixed the reading as (c) before the arms were spawned, so S = 16.1.0 / gap 16.2.0 here and on all four Fable 5.1 arms, which all show the same divergence. ============================================================================== RUN next.js--claude-fable-5-1--v4-b--2026-09-05 What Claude Fable 5.1 gets right about next.js — battery v4-b, tested 2026-09-05 URL: https://stalepriors.com/runs/next.js--claude-fable-5-1--v4-b--2026-09-05 JSON: https://stalepriors.com/data/next.js/fable-5-1-v4-b.json Library: next.js 16.3.4 (npm), verified 2026-09-05 Model: Claude Fable 5.1 (claude-fable-5-1), stated cutoff 2026-06 Version attribution stops at: 16.1.0 (2025-12-18), lag ~6 months Oldest release it could not place: 16.2.0 (2026-03-18) Battery: next.js/v4-b, 0 tasks, tool uses during test: 0 Tested: 2026-09-05 Summary: Blind replicate of `v4-a`: same stored file, same order (boundary question first). It reached the same self-placement (16.1 describable, 16.2 the first bare number) and the same five correct rungs, and its 16.1.0 answer is the one that is CORRECT under both grading readings - it names the file-system caching headline and binds it to 16.1 without disclaiming the attribution. D = 16.1.0, S = 16.1.0, gap 0. Poison and ceiling controls clean. --- non-finding [correct] next 16.0.0: LADDER GRADE - CORRECT. The fullest 16.0.0 answer in the battery: Turbopack default, Cache Components as the successor to `experimental.dynamicIO` with `cacheComponents: true`, the `use cache` directive, `cacheLife()`/`cacheTag()`, `updateTag()`/`refresh()`, `revalidateTag` taking a cache profile, `proxy.ts` exporting `proxy`, stable `reactCompiler`, the DevTools MCP server, React 19.2, the async-only request APIs, the removal of `next lint` and AMP, Node 20.9, layout deduplication and incremental prefetching, the adapters API and `experimental.turbopackFileSystemCacheForDev`. Every one of those is in the release post. --- non-finding [correct] next 16.1.0: LADDER GRADE - CORRECT, and the arm that separates the two readings. It named Turbopack file-system caching for development promoted toward stable, reducing repeated compile times across restarts, and bound it to 16.1.0 - the release's headline change - and its hedge is about further content ("I cannot name additional config keys or file conventions with confidence beyond that"), not about that attribution. HARNESS.md clause 2 grades that CORRECT under both readings. D = 16.1.0. --- non-finding [correct] next 15.3.0: LADDER GRADE - CORRECT. Turbopack build alpha, `next-rspack`, `instrumentation-client.js|ts` described as running before hydration, `onNavigate` with `preventDefault`, `useLinkStatus`, TypeScript plugin performance. --- non-finding [correct] next 15.1.0: LADDER GRADE - CORRECT. React 19 stable for both routers, `after()` promoted from `unstable_after` with the correct import, `forbidden()`/`unauthorized()` behind `experimental.authInterrupts`, collapsed ignore-listed stack frames. --- non-finding [correct] next 15.5.0: LADDER GRADE - CORRECT. Turbopack builds beta, Node.js middleware stable, `typedRoutes` stable, route export validation, `next typegen`, `next lint` deprecated, and the four 16 deprecation warnings. --- non-finding [imprecision] next 16.0.0: IMPRECISION inside a CORRECT rung. The arm listed `next build --debug-prerender` among 16.0.0's changes and said `legacyBehavior` on `next/link` was removed in 16.0.0. The 16 release post lists neither: `legacyBehavior` is a 15.5 deprecation warning and the 16 removals table does not name it. Does not void the grade - a rung is CORRECT on naming an anchor - but it is wrong content behind a correct answer and is recorded as such. --- non-finding [correct] next 15.7.0: POISON RUNG. 15.7.0 has never been published in any form - zero versions of any kind match 15.7.* on the npm registry, checked this session. "I believe this version was never released... I have no knowledge of a 15.6 or 15.7 minor." Correct on both. P4 holds on this arm, so its CORRECT grades stand. --- non-finding [correct] next 16.3.0: CEILING RUNG. 16.3.0 (2026-08-03) is above this subject's stated cutoff. "Cannot describe. I have no content for this version and cannot confirm whether it has been released." P5 holds on this arm. --- non-finding [context] stated cutoff: INSTRUMENT OBSERVATION. "My system context states a cutoff of June 2026." Stated independently of its twin `v4-a`, which stated the same. Both are Claude Fable 5.1, not the Claude Fable 5 of the Index's earlier Fable runs. ============================================================================== RUN next.js--claude-fable-5-1--v4-c--2026-09-05 What Claude Fable 5.1 gets right about next.js — battery v4-c, tested 2026-09-05 URL: https://stalepriors.com/runs/next.js--claude-fable-5-1--v4-c--2026-09-05 JSON: https://stalepriors.com/data/next.js/fable-5-1-v4-c.json Library: next.js 16.3.4 (npm), verified 2026-09-05 Model: Claude Fable 5.1 (claude-fable-5-1), stated cutoff 2026-06 Version attribution stops at: 16.1.0 (2025-12-18), lag ~6 months Oldest release it could not place: 16.2.0 (2026-03-18) Battery: next.js/v4-c, 0 tasks, tool uses during test: 0 Tested: 2026-09-05 Summary: Ladder first, boundary question last - the position every published battery in the Index uses. It produced the same five correct rungs and the same self-placement as the two arms that were asked the boundary question first: 16.1 describable, 16.2 the first bare number. D = 16.0.0 strict / 16.1.0 lenient, S = 16.1.0. Poison and ceiling controls clean. --- non-finding [correct] next 16.0.0: LADDER GRADE - CORRECT. Cache Components with `cacheComponents: true` and the `use cache` directive, PPR folded in and `experimental.ppr` replaced, `cacheLife()`/`cacheTag()`, `updateTag()`/`refresh()`, Turbopack default with a `--webpack` opt-out, `experimental.turbopackFileSystemCacheForDev`, `proxy.ts`, React 19.2, stable `reactCompiler`, layout deduplication, incremental prefetching, DevTools MCP, the adapters API and the removals. Attributed to 16.0.0 and supported by the release post. --- non-finding [context] next 16.1.0: LADDER GRADE - DISPUTED, the same shape as `v4-a` and disclosed the same way. It named Turbopack file-system caching moving toward stable for 16.1 - the correct headline - and then wrote "I may be blending in 16.0.x patch-release notes", which disclaims the attribution. ABSTAIN under HARNESS.md clause 2 (D = 16.0.0); CORRECT under the wrapper-hedge reading (D = 16.1.0). Both recorded. That the disputed shape falls on one `-sc` arm and one `-cs` arm is why the battery's comparison survives the ambiguity. --- non-finding [correct] next 15.3.0: LADDER GRADE - CORRECT. Turbopack build alpha, `next-rspack`, `instrumentation-client.js|ts` at the project root, `onNavigate` and `useLinkStatus()`, TypeScript plugin performance. --- non-finding [correct] next 15.1.0: LADDER GRADE - CORRECT. React 19 stable, `after()` promoted from `unstable_after`, `forbidden()`/`unauthorized()` behind `experimental.authInterrupts`, improved source maps and error overlay. --- non-finding [correct] next 15.5.0: LADDER GRADE - CORRECT, and one of the two arms to name all three route props helpers. `typedRoutes` moved out of experimental, route export validation, `PageProps`/`LayoutProps`/`RouteContext`, `next typegen`, plus the 16 deprecation warnings. --- non-finding [correct] next 15.7.0: POISON RUNG. 15.7.0 has never been published in any form - zero versions of any kind match 15.7.* on the npm registry, checked this session. "I believe this was never released as a minor. The last 15.x minor I know of is 15.5; after 16.0 the 15 line continued only as 15.5.x patch releases." Correct. P4 holds on this arm, so its CORRECT grades stand. --- non-finding [correct] next 16.3.0: CEILING RUNG. 16.3.0 (2026-08-03) is above this subject's stated cutoff. "Cannot describe. I have no content attached to this version and cannot confirm it exists." P5 holds on this arm. --- non-finding [context] stated cutoff: INSTRUMENT OBSERVATION. "My stated training cutoff is June 2026." Third of four Fable 5.1 draws to state it, in the opposite prompt order from `v4-a`/`v4-b`, so the value does not depend on where the boundary question sits. ============================================================================== RUN next.js--claude-fable-5-1--v4-d--2026-09-05 What Claude Fable 5.1 gets right about next.js — battery v4-d, tested 2026-09-05 URL: https://stalepriors.com/runs/next.js--claude-fable-5-1--v4-d--2026-09-05 JSON: https://stalepriors.com/data/next.js/fable-5-1-v4-d.json Library: next.js 16.3.4 (npm), verified 2026-09-05 Model: Claude Fable 5.1 (claude-fable-5-1), stated cutoff 2026-06 Version attribution stops at: 16.1.0 (2025-12-18), lag ~6 months Oldest release it could not place: 16.2.0 (2026-03-18) Battery: next.js/v4-d, 0 tasks, tool uses during test: 0 Tested: 2026-09-05 Summary: Blind replicate of `v4-c`. Same five correct rungs, same self-placement, and - like `v4-b` in the opposite order - a 16.1.0 answer that is CORRECT under both grading readings. The four Fable 5.1 arms split two-and-two on the 16.1.0 grade, and the split runs across the prompt orders rather than with them: one disputed and one clean arm in each cell. D = 16.1.0, S = 16.1.0, gap 0. --- non-finding [correct] next 16.0.0: LADDER GRADE - CORRECT, and the most precisely attributed of the four. It named `revalidateTag(tag, profile)` with 'max' as an example profile, `experimental.adapterPath` for the adapters API, the TypeScript 5.1 floor alongside Node 20.9, `images.imageSizes` dropping 16, and `serverRuntimeConfig`/`publicRuntimeConfig` removal - each of which is in the 16 release post's own tables. --- non-finding [correct] next 16.1.0: LADDER GRADE - CORRECT under both readings. Turbopack file-system caching moved toward or into stable for dev, bound to 16.1.0; the hedge that follows is about completeness ("treat the above as an outline, not a feature list"), not about the binding. D = 16.1.0. --- non-finding [correct] next 15.3.0: LADDER GRADE - CORRECT. `next build --turbopack` alpha, `next-rspack`, `instrumentation-client.js|ts`, `onNavigate`, `useLinkStatus()`, TypeScript plugin performance. --- non-finding [correct] next 15.1.0: LADDER GRADE - CORRECT. React 19 stable, `after()` stable, `forbidden()`/`unauthorized()` behind `experimental.authInterrupts` with both file conventions, error-debugging improvements. --- non-finding [correct] next 15.5.0: LADDER GRADE - CORRECT. Turbopack builds beta, Node.js middleware stable, `typedRoutes` stable, route export validation, `next typegen`, the three route props helpers, `next lint` deprecated, and the 16 deprecation warnings including `images.domains`. --- non-finding [correct] next 15.7.0: POISON RUNG. 15.7.0 has never been published in any form - zero versions of any kind match 15.7.* on the npm registry, checked this session. "I have no knowledge of this version... I would say it was not released, with the caveat that I'd be describing absence of knowledge rather than certainty" - the most carefully qualified of the six poison answers, and it still declines. P4 holds on this arm, so its CORRECT grades stand. --- non-finding [correct] next 16.3.0: CEILING RUNG. 16.3.0 (2026-08-03) is above this subject's stated cutoff. "cannot describe. I don't know whether it exists." P5 holds on this arm. --- non-finding [context] stated cutoff: INSTRUMENT OBSERVATION. "I'm told my training cutoff is June 2026." Fourth of four Fable 5.1 draws stating it. The value is stable across both prompt orders and both replicate pairs; what changed is the model behind the alias, not the draw. --- non-finding [context] self-placement, question (a): The clearest statement of the (a)/(c) divergence in the battery: "The latest version I know of at all is 16.2... The most recent release whose contents I can actually describe is 16.1 (approx. December 2025), and only at outline level; the most recent one I can describe in detail is 16.0." Three tiers - describable in detail, describable in outline, bare number - where the Index's schema has two. ============================================================================== RUN next.js--claude-fable-5--v2-d--2026-09-02 What Claude Fable 5 gets right about next.js — battery v2-d, tested 2026-09-02 URL: https://stalepriors.com/runs/next.js--claude-fable-5--v2-d--2026-09-02 JSON: https://stalepriors.com/data/next.js/fable-5-v2-d.json Library: next.js 16.3.4 (npm), verified 2026-09-02 Model: Claude Fable 5 (claude-fable-5), stated cutoff 2026-01 Version attribution stops at: 16.0.0 (2025-10-22), lag ~3 months Oldest release it could not place: 16.1.0 (2025-12-18) Battery: next.js/v2-d, 9 tasks, tool uses during test: 0 Tested: 2026-09-02 Summary: The second below-floor control arm, and the more informative of the two: its next.js attribution stops at 16.0.0, one minor below the probe band rather than sixteen releases below it, and it passed the battery's internal control cleanly. So its answers read as evidence about the probes rather than about the subject. It struck the same two probes DERIVABLE as the other control, and it produced the sharpest instance in the battery of a category the Index has been tracking since valibot: asked about the image disk-cache budget, it stated flatly that no such rule exists — 'if your question presumes a percentage-of-the-volume rule, I believe that premise is false' — and correctly identified the 50 MB figure a reader might be thinking of as cacheMaxMemorySize, an unrelated in-memory cache. Confidently right about what a thing is not, and wrong that anything replaced it. --- non-finding [context] error.tsx retry prop: DERIVABLE — task 1(b). 'It is purely a client-side state reset on Next's internal error boundary (setState({ error: null }) + a transition) ... No network request is made, the Server Component does not re-render, and the database query does not run again.' Correct in mechanism and in what does re-execute, from below the band. Prediction P4 confirmed by both controls independently. --- non-finding [context] catchError (next/error): DERIVABLE — task 2(b). Named the NEXT_REDIRECT digest, the RedirectBoundary, unstable_rethrow, and the full-document-request caveat. Prediction P5 confirmed. --- non-finding [miss] images.maximumDiskCacheSize: Task 5(a) — the sharpest single answer in the battery, and a third instance of 'right about the removal, wrong about the replacement'. It denied the premise: 'To the best of my knowledge, there is no default size cap at all ... it does not compute any fraction of the volume at startup or any other time ... I'm reasonably confident no such rule exists in stable Next.js.' It then pre-empted the likely confusion correctly — 'the 50 MB default cap you may be thinking of is cacheMaxMemorySize, which is the in-memory ISR/data cache, unrelated to images' — and declined to invent a config key for (b), which is the right behaviour. Everything it ruled out, it ruled out correctly. The one thing it did not consider is that the cap shipped, in a patch, four months above its boundary. Not chargeable: 16.1.7 is 2026-03-16, two months above this subject's stated cutoff. --- non-finding [miss] ImageResponse default font: Task 3: 'Noto Sans ... Confidence: high on the family name.' The fourth of four draws to name it. Not chargeable — 16.2.0 is above this subject's cutoff. --- non-finding [miss] Link transitionTypes: Task 4: 'transitionTypes is not a real next/link prop in any stable Next.js I know of', with the React unknown-attribute warning predicted for both routers. Not chargeable — above cutoff. --- non-finding [miss] experimental.clientSegmentCache: Task 8: experimental.clientSegmentCache: false, at moderate confidence, with the mechanism described correctly and the Next-Router-Segment-Prefetch header named. The key was removed in 16.0.3 on 2025-11-13, two months inside this subject's stated cutoff, so the miss is chargeable in principle and is not charged only because this is a control arm. Both controls made this recommendation; it wants a charging battery. --- non-finding [correct] images.qualities: Task 7, the internal control, passed cleanly: the [75] default, the 400 enforcement asserted with the clamp named as the alternative, and the release dated correctly to 16.0.0 in October 2025 with the 15.5 deprecation warning that preceded it. This subject can place a change it holds, in a major, to the release. That is what makes its failures above the band readable as knowledge gaps rather than dating failures — and it is the same shape as the two Opus 5 draws. --- non-finding [context]: Task 9: it declined to guess a version for a rule it had denied — 'Since I claim the percentage-of-the-volume rule doesn't exist, there is no release to name for it; if it does exist, it postdates my reliable knowledge and this is a stated gap, not a guess at a version.' The most honest attribution answer of the four, and the only one that separates 'I do not know the date' from 'I do not know the fact'. --- non-finding [context]: Boundary reproduced exactly: next.js/v1 measured this subject at 16.0.0 / 16.1.0 on 2026-08-31 with a different prompt, and v2 lands on the same pair. All three subjects reproduced their v1 next.js boundary under v2. Not re-charged; the recency finding belongs to v1. ============================================================================== RUN next.js--claude-fable-5--v3-c--2026-09-02 What Claude Fable 5 gets wrong about next.js — battery v3-c, tested 2026-09-02 URL: https://stalepriors.com/runs/next.js--claude-fable-5--v3-c--2026-09-02 JSON: https://stalepriors.com/data/next.js/fable-5-v3-c.json Library: next.js 16.3.4 (npm), verified 2026-09-02 Model: Claude Fable 5 (claude-fable-5), stated cutoff 2026-01 Version attribution stops at: 16.0.0 (2025-10-22), lag ~3 months Oldest release it could not place: 16.1.0 (2025-12-18) Battery: next.js/v3-c, 5 tasks, tool uses during test: 0 Tested: 2026-09-02 Summary: The Fable 5 half of next.js/v3, run 2026-09-02 after both Fable arms of the original battery died on an API-side safeguard error without emitting a token. The stored prompt was re-sent verbatim; nothing was reworded. This is the charging arm and it charges one finding: asked in one word whether per-part prefetching can be turned off, it answered "Yes" and wrote experimental: { clientSegmentCache: false }, a key removed in patch 16.0.3 on 2025-11-13, two months inside its own stated cutoff. Handed the same key back in task 2 it said the build accepts the config, and in task 5(a) it committed to the key still being recognised on current stable. What separates this draw from the Sonnet 5 draw that failed the same probe is that it did not assert a mechanism it lacked: it put the whole answer at about 60 percent, named the opposite branch explicitly, and reproduced the exact warning string the build would print if the key were gone - which it is. The finding is charged on the artefact under the code-vs-claim rule and the hedge is recorded with it. On the internal control it passed cleanly, naming proxy.ts exporting proxy and dating the rename to 16.0.0, which is right and is a major; the same draw could not reach a patch for the removal. Its twin v3-d, blind to this one, answered task 1 "No" and got the entire surface right, which falsifies the battery's first prediction. Both Fable draws read the same boundary this battery read in v1 and v2 - 16.0.0 describable, 16.1.0 not - so on this subject the boundary did not move between batteries or between twins. --- F1 [S2 silently-wrong] Answers "yes" and configures experimental.clientSegmentCache, a key removed in patch 16.0.3 API: experimental.clientSegmentCache Changed in next.js 16.0.3 (2025-11-13), kind: removed Chargeability: Removed 2025-11-13, two months before this draw's stated cutoff of January 2026. Not back-filled from another run: the twin v3-d states the same cutoff independently and charges nothing, because it got the surface right. Model belief: Task 1(a), one word: "Yes." Task 1(b): "It is controlled by the clientSegmentCache flag ... Disabling it reverts to the older whole-route prefetch - one RSC payload request per link". Task 2(a): "the key is still recognised, so the build accepts the config". Task 2(b): "With false you get the legacy behaviour - a single prefetch per link carrying the route's RSC payload, duplicating shared-layout data across links". Task 5(a): "I believe experimental.clientSegmentCache is still recognised on current stable (it flipped default rather than disappearing)". Wrong: experimental: { clientSegmentCache: false } Correct: experimental: { prefetchInlining: true } Impact: The team asked how to collapse a prefetch burst into one request per link and gets a config file that builds, deploys and does nothing. An unrecognised key under experimental warns and is dropped rather than failing the build, so the file reads as applied and the burst continues unchanged. The reader is also told the reverted behaviour will duplicate shared-layout data across links, which is a specific and false description of what their app will now do. Scope note: The hedge is part of the record and is not deducted from the charge. This draw put the answer at "moderate (~60%)", wrote out the branch in which it is wrong, and reproduced the warning the build actually prints - "Invalid next.config.ts options detected: Unrecognized key(s) in object: 'clientSegmentCache' at experimental" - as the counterfactual. It is charged because the task demanded a one-word verdict before any explanation, the verdict was "Yes", and the configuration written under it is inert. The code-vs-claim rule (prompts/nextjs.md v1) scores the artefact a reader would copy. Compared against next.js/v3-a, which failed the same probe while asserting the key "is a recognized key under experimental in the config schema Next.js validates against", this is the same wrong answer delivered with the uncertainty intact - the same severity, because the failure mode of a silently inert config does not depend on how confident its author was. Source: https://unpkg.com/next@16.0.2/dist/server/config-shared.d.ts (2025-11-12) — "clientSegmentCache?: boolean | 'client-only';" Source: https://unpkg.com/next@16.0.3/dist/server/config-shared.d.ts (2025-11-13) — "(no occurrence of clientSegmentCache in the file)" Source: https://unpkg.com/next@16.3.4/dist/server/config.js (2026-08-31) — "if (issue.code === 'unrecognized_keys' && issue.path[0] === 'experimental') {" --- non-finding [correct] proxy.ts / export function proxy: Task 4, the internal control, passed cleanly. Named proxy.ts at the project root with the named export proxy, wrote a correct matcher-scoped redirect, named middleware.ts / export function middleware as the older form, and said correctly that it still works on 16 and warns. It also correctly identified the Next 12 per-directory _middleware.ts as a still older form that was removed in 12.2. Under JOURNAL/030(b) this is what makes the run's attribution answers readable at all: the subject holds the change, so its dating of it measures attribution rather than absence. --- non-finding [miss] images.maximumResponseBody: Task 3, and the reason it is not charged against this subject. Answered that the browser receives a resized AVIF or WebP with HTTP 200 and that no documented byte-size cap on the source file exists in self-hosted Next.js. The truth is a fixed 50 MB body cap that throws a 413 while streaming. images.maximumResponseBody arrived in 16.1.5 on 2026-01-26, the same month as this subject's stated cutoff, so it sits outside the fairness window under the same-month rule and is not a chargeable miss - the battery fixed this reading in advance for v3-a and v3-c both. Two details worth keeping: the draw declined to invent a key or a number, saying so explicitly - "I cannot honestly write a 5 MB config, and I won't invent one" - and it reached instead for sharp's 268 megapixel input limit and the keys experimental.imgOptMaxInputPixels and experimental.imgOptTimeoutInSeconds. Those two key names were not verified this session and the Index makes no claim about them either way. --- non-finding [correct] images.maximumResponseBody: Task 3(c), the fixed-policy versus resource discriminator, answered correctly and for the correct reason despite the wrong answer in (a) and (b): "No. Neither limit is resource-adaptive: the pixel cap is a decompression-bomb guard and the timeout is fixed config, so more RAM/disk changes nothing about the defined behaviour." It then separated an undefined failure - the process being OOM-killed mid-decode - from a framework limit. This is the sub-question next.js/v3-e and v3-f landed on opposite sides of; both Fable draws land on the correct side. --- non-finding [context] experimental.clientSegmentCache: Attribution, and the patch result holds for a fourth draw. This draw dates the middleware-to-proxy rename to 16.0.0, ~October 21 2025, which is right to the release and is a major it demonstrably holds. Asked in the same task for the release at which clientSegmentCache stopped being recognised, it committed to the key still being recognised and named 16.0.0 as the fallback if it were wrong. The removal is 16.0.3. Across two subjects and three batteries no draw has yet attributed any change to a patch release, and this draw shows the split inside one task: the major it holds is dated correctly, the patch is unreachable even as a hypothetical. --- non-finding [context]: The boundary did not move for this subject, which is the contrast the battery could not draw from Sonnet 5. Fable 5 x next.js now has four readings across three batteries - v1, v2-d, and both v3 twins - and all four are identical: 16.1.x known by name, 16.0.0 (2025-10-22) the most recent release whose contents it can describe, nothing describable at 16.1.0 (2025-12-18). Sonnet 5 on the same library reads 15.0.0, 15.0.0, 15.3.0 and 15.5.0 across the same three batteries. JOURNAL/033 left open whether v3 moved Sonnet's boundary or whether v1 and v2 happened to agree; this run does not settle that, but it does show the instrument holding perfectly still on a second subject over the same three batteries, so whatever moves Sonnet's reading is not something every battery does to every subject. --- non-finding [context]: The arm ran. Both Fable arms of this battery failed on 2026-09-02 with an API-side safeguard error (invalid_request, [reasoning_extraction]) before emitting a token, were retried once each and failed identically, and were recorded in JOURNAL/033 as void rather than as runs. The stored prompt at prompts/sent/nextjs-v3.txt was re-sent byte-identical in a later session and both arms completed without incident. The earlier failure is therefore transient and infrastructural, and is not a property of this prompt or this subject. No wording changed between the void attempt and this run. ============================================================================== RUN next.js--claude-fable-5--v3-d--2026-09-02 What Claude Fable 5 gets right about next.js — battery v3-d, tested 2026-09-02 URL: https://stalepriors.com/runs/next.js--claude-fable-5--v3-d--2026-09-02 JSON: https://stalepriors.com/data/next.js/fable-5-v3-d.json Library: next.js 16.3.4 (npm), verified 2026-09-02 Model: Claude Fable 5 (claude-fable-5), stated cutoff 2026-01 Version attribution stops at: 16.0.0 (2025-10-22), lag ~3 months Oldest release it could not place: 16.1.0 (2025-12-18) Battery: next.js/v3-d, 5 tasks, tool uses during test: 0 Tested: 2026-09-02 Summary: The second Fable 5 draw of next.js/v3, blind to v3-c, run from the same stored prompt in the same session. It is the -b twin by role and charges nothing, and it got the battery's target surface entirely right. Task 1(a), one word: "No." It then gave the correct options - , manual router.prefetch, flattening the route tree - and noted correctly that router.prefetch also goes through the segment cache and so buys timing rather than request count. Task 2(a): warns and continues, with the invalid-config warning reproduced close to verbatim. Task 2(b): "Nothing. The key is ignored", for the right reason. It named clientSegmentCache only to place it on 15.x as an opt-in that 16 removed. That falsifies prediction P1, which expected all four Sonnet 5 and Fable 5 draws to reach the dead key, and it makes the fourth battery running in which the non-charging twin holds the better answer. Everything it got wrong is either outside its fairness window (the 50 MB image body cap, 16.1.5, same month as its stated cutoff) or is the patch-granularity result again: it placed the clientSegmentCache removal at 16.0.0, a major, where it shipped in 16.0.3 three weeks later, while dating the proxy.ts rename to 16.0.0 correctly. --- non-finding [correct] experimental.clientSegmentCache: Tasks 1 and 2, the battery's whole target surface, answered correctly in both directions. Task 1(a): "No." with the reasoning that per-segment prefetching is the default and only routing implementation on 16 and the old whole-route prefetch mode went with the legacy router. Task 1(b) gave real options - paired with a manual router.prefetch on hover, flattening nested layouts, or staying on 15.x - and correctly qualified the manual route: "router.prefetch also goes through the segment cache, so it issues the same per-part requests; it only lets you control when, not how many." Task 2(a): warns and continues, with the warning text reproduced close to verbatim, and the correct general rule that config validation warns on unknown keys rather than failing. Task 2(b): "Nothing. The key is ignored." It reached the key only to place it correctly on 15.x as an opt-in and to say 16 removed it. The one thing it could not do is date the removal to the patch. --- non-finding [correct] proxy.ts / export function proxy: Task 4, the internal control, passed. proxy.ts at the project root or src/, named export proxy, a correct matcher-scoped redirect that also preserves the origin path as a query parameter; middleware.ts / export function middleware named as the older form, still working on 16 and deprecated with a rename notice; the Next 12 pages/**/_middleware.ts convention correctly named as removed in 12.2. It put the named export at about 85 percent confidence and the rename itself at high confidence, which matches the outcome. --- non-finding [miss] images.maximumResponseBody: Task 3, outside this subject's fairness window. Answered that self-hosted Next.js enforces no documented byte-size cap on the upstream source image, that the server downloads all 80 MB and the browser receives an optimized image with HTTP 200, and consequently that there is no key it could write to lower the limit to 5 MB. The truth is images.maximumResponseBody, defaulting to 50000000, enforced while streaming with a 413. It arrived in 16.1.5 on 2026-01-26, the same month as this subject's stated cutoff, so the same-month rule parks it: not a chargeable miss. Recorded because the draw did something the scale gives it no credit for - it stated the possibility that a cap exists and that it was failing to recall it, at "roughly 50/50", and declined to name a key rather than inventing one. --- non-finding [correct] images.maximumResponseBody: Task 3(c) answered correctly and for the correct reason: "No. Whatever happens is the same on a beefy server, because the outcome is determined by framework policy plus sharp's pixel limit, not by available RAM/disk." Like its twin it separated an undefined failure - the process not surviving the decode - from a defined framework limit. Both Fable draws land on the correct side of the discriminator that split the two Opus draws in v3-e and v3-f. --- non-finding [context] experimental.clientSegmentCache: Attribution, and the patch result again. Task 5(a): the removal placed at 16.0.0, October 21 2025, "when the rewritten router/segment cache became unconditional", with the draw explicitly flagging that the specific release was a guess at about 60 percent while the 16.0.0 date itself was not. It shipped in 16.0.3, three weeks later. Task 5(c): the middleware-to-proxy rename placed at 16.0.0, correct. The same shape as v3-c and as both Opus draws before it - the major is reached, the patch is not, inside a single task, by a subject that demonstrably holds both changes. Prediction P2 is now confirmed on six scored draws of six. --- non-finding [context]: The twins agree about the boundary and disagree about the surface, which is the reverse of what the replication rule was written to catch. JOURNAL/023 introduced replication because the self-report is the unstable half and the code is the stable half - langchain/v1r got two byte-identical draws 399 days apart on the boundary while both wrote the same stale imports. Here both Fable draws report an identical boundary (16.0.0 describable, 16.1.0 not) and an identical stated cutoff, and split on the code: v3-c writes the dead key, v3-d refuses it. Recorded as a counterexample rather than a correction - one battery does not overturn the langchain result - but it is the first pair in the Index where the instrument held still and the answer moved. ============================================================================== RUN next.js--claude-fable-5--v1--2026-08-31 What Claude Fable 5 gets wrong about next.js — battery v1, tested 2026-08-31 URL: https://stalepriors.com/runs/next.js--claude-fable-5--v1--2026-08-31 JSON: https://stalepriors.com/data/next.js/fable-5.json Library: next.js 16.3.3 (npm), verified 2026-08-31 Model: Claude Fable 5 (claude-fable-5), stated cutoff 2026-01 Version attribution stops at: 16.0.0 (2025-10-22), lag ~3 months Oldest release it could not place: 16.1.0 (2025-12-18) Battery: next.js/v1, 10 tasks, tool uses during test: 0 Tested: 2026-08-31 Summary: The strongest result recorded in this dataset. Fable 5 answers nine of ten tasks as a model that has actually read Next.js 16: proxy.ts, cacheComponents, unprefixed cacheLife/cacheTag, the eslint script, Turbopack by default, images.qualities [75], the four-hour image TTL correctly attributed to 16. It knows 16.1 exists and shipped mid-December 2025 — within two weeks of its own stated cutoff — and correctly declines to describe its contents. Two findings survive: it still generates the deprecated single-argument revalidateTag, and it tells you that parallel-route default.js files are optional when 16.0.0 made builds fail without them. Its stated cutoff is 2026-01, four months earlier than Opus 5's, and it is the more current model on this library. That dissociation, first recorded on zod, reproduces here on a different library and a different kind of change. --- F1 [S1 breaks-build] Single-argument revalidateTag in a generated Server Action API: revalidateTag Changed in next.js 16.0.0 (2025-10-22), kind: behavior-changed Model belief: Wrote `revalidateTag('products')` in the action, while stating correctly in the surrounding prose that "in Next 16 with Cache Components enabled ... revalidateTag accepts a cache-life profile as a second argument" and that `updateTag` gives read-your-writes. At question (c) it gives the right answer: "In Next 16 ... it accepts up to two: revalidateTag(tag, profile)." Wrong: 'use server' import { revalidateTag } from 'next/cache' export async function updateProduct(id: string, formData: FormData) { await db.product.update({ where: { id }, data: { ... } }) revalidateTag('products') } Correct: 'use server' import { revalidateTag, updateTag } from 'next/cache' export async function updateProduct(id: string, formData: FormData) { await db.product.update({ where: { id }, data: { ... } }) revalidateTag('products', 'max') // stale-while-revalidate // or updateTag('products') for read-your-writes inside the action } Impact: TypeScript project: the single-argument form is documented as producing a TypeScript error, and `next build` type-checks by default. The model knows the two-argument form and does not use it — the same knowing-but-not-applying failure Opus 5 shows on the same line. All three subjects generated this identical call. It is the most reliable finding in the battery, and the one worth putting at the top of the correction pack. Source: https://nextjs.org/docs/app/guides/upgrading/version-16 — "revalidateTag now requires a second argument specifying a cacheLife profile. The single-argument form is deprecated and will produce a TypeScript error." Source: https://nextjs.org/blog/next-16 (2025-10-21) — "`revalidateTag()` signature | Now requires `cacheLife` profile as second argument for stale-while-revalidate behavior" --- F2 [S2 silently-wrong] States a parallel route builds without default.js; on 16.0.0 the build fails API: parallel routes default.js Changed in next.js 16.0.0 (2025-10-22), kind: stricter Model belief: "Strictly, the route builds without default.tsx, but the moment you add child routes under /dashboard, any route the slot can't match hard-404s on reload — so treat default.tsx as required in practice." Wrong: // The claim, not the code: default.tsx presented as a soft requirement // whose failure mode is a runtime 404 on reload. Correct: // app/dashboard/@panel/default.tsx — required for the build to succeed export default function PanelDefault() { return null } Impact: 16.0.0 made explicit `default.js` files mandatory for every parallel-route slot, and builds fail without them. The consequence is inverted: it is a build-time failure, not a runtime 404. A developer who accepts "strictly, it builds without it" and trims the file gets a red CI run they were told to expect as a browser 404. The model's own file listing does include default.tsx, so the generated code is correct — only the claim about it is wrong, which is why this is S2 and not S1. Source: https://nextjs.org/blog/next-16 (2025-10-21) — "Parallel routes `default.js` | All parallel route slots now require explicit `default.js` files; builds fail without them. Create `default.js` that calls `notFound()` or returns `null` for previous behavior" --- non-finding [correct] proxy.ts: Task 1: proxy.ts with a default-exported proxy function, correctly dated to Next.js 16 and correctly described as a rename with middleware.ts deprecated but functional. --- non-finding [correct] cacheComponents: Task 4: `cacheComponents: true` at the top level of next.config.ts, explicitly identified as subsuming the old experimental.ppr flag — the only subject to get this right. Both other subjects shipped the removed experimental.ppr configuration. --- non-finding [correct] cacheLife / cacheTag: Task 5: `'use cache'` with unprefixed `cacheLife` / `cacheTag`, with the `unstable_cache` form given as the pre-16 fallback. --- non-finding [correct] next lint / next build: Task 6: "next lint was removed in Next 16 (deprecated in 15.5)" — the correct dating, which Opus 5 also gets right — plus Turbopack as the default bundler for dev and build with `--webpack` as the opt-out. --- non-finding [correct] images.minimumCacheTTL: Task 10: "The default changed in Next 16: 4 hours (14400 seconds). On Next ≤15 the default was 60 seconds." Correct value and correct release — the only subject to get both. --- non-finding [correct]: Task 3: async params in the OG image route. Task 7: remotePatterns. Task 8: knew `images.qualities` defaults to `[75]` in 16 and that quality 90 is not served. --- non-finding [correct]: Version recency: version attribution stops at 16.0.0 (2025-10-22) against a stated 2026-01 cutoff, and it names 16.1 as existing and shipping "around mid-December 2025" — 16.1.0 shipped 2025-12-18, two weeks before the cutoff — while correctly refusing to describe its contents. No version-recency finding is charged against this subject. It is the only run in the dataset so far where the model's version attribution reaches its own cutoff. --- non-finding [imprecision] images.qualities: Task 8: claims the optimizer "rejects qualities not in the allowlist with a 400 Bad Request," self-rated "slightly less than certain it 400s rather than clamping." The upgrade guide says the quality is coerced to the nearest allowed value. Hedged, and paired with the correct fix (add 90 to images.qualities), so not charged under the code-vs-claim rule. --- non-finding [imprecision] images.domains: Task 7: says `images.domains` is "deprecated and removed in Next 16." It is deprecated in 16.0.0, not removed. No consequence — the model used remotePatterns. ============================================================================== RUN next.js--claude-opus-5--v2-a--2026-09-02 What Claude Opus 5 gets wrong about next.js — battery v2-a, tested 2026-09-02 URL: https://stalepriors.com/runs/next.js--claude-opus-5--v2-a--2026-09-02 JSON: https://stalepriors.com/data/next.js/opus-5-v2-a.json Library: next.js 16.3.4 (npm), verified 2026-09-02 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 16.0.0 (2025-10-22), lag ~7 months Oldest release it could not place: 16.1.0 (2025-12-18) Battery: next.js/v2-a, 9 tasks, tool uses during test: 0 Tested: 2026-09-02 Summary: The test arm of the battery that generalised the capability probe to a third library. Four findings charged: one S1 (a guessed image-config key, which exits the build rather than warning), and three S2 — a config key removed in a patch and now silently inert, the OG-image default typeface, and a Link prop the draw asserted does not exist. The battery's two purest capability probes charged nothing, and that is the methodological result: asked how to re-run a failed server render and how to scope an error boundary to one widget, the draw hand-rolled both, and both hand-rolled answers work. The Index's severity scale has no slot for 'correct code a framework API now supersedes', so those are recorded as misses without an S-level rather than inflated into findings. Two probes were struck DERIVABLE by the control arm and one was voided by an ambiguity in its own wording. --- F1 [S1 breaks-build] Invents images.maximumCacheSize, and an unrecognised images key exits the build API: images.maximumCacheSize Changed in next.js 16.1.7 (2026-03-16), kind: added Chargeability: 16.1.7 published 2026-03-16, two months before the stated 2026-05 cutoff. Model belief: "I'm committing to `images.maximumCacheSize` as the key at roughly 50% confidence — verify against the next.config typings before relying on it, since an unknown images.* key is a hard config validation error at boot, so you'll find out immediately." The draw then wrote that key into both the 1 GB cap and the disable-caching config. Wrong: images: { remotePatterns: [...], maximumCacheSize: 1024 * 1024 * 1024, } Correct: images: { remotePatterns: [...], maximumDiskCacheSize: 1_000_000_000, // 0 disables the disk cache } Impact: The config does not merely fail to take effect. `normalizeNextConfigZodErrors` sets `shouldExit` for any validation issue whose path starts at `images`, so the build exits. The draw named this consequence itself and shipped the key anyway, which is the code-vs-claim rule's exact case: the artefact is the config, and the config does not build. Source: https://unpkg.com/next@16.3.4/dist/shared/lib/image-config.js (2026-08-31) — "maximumDiskCacheSize: undefined," Source: https://unpkg.com/next@16.3.4/dist/server/config.js (2026-08-31) — "if (issue.path[0] === 'images') { // We exit the build when encountering an error in the images config shouldExit = true; }" --- F2 [S2 silently-wrong] Recommends experimental.clientSegmentCache, removed in 16.0.3 and now silently ignored API: experimental.clientSegmentCache Changed in next.js 16.0.3 (2025-11-13), kind: removed Chargeability: Removed 2025-11-13, six months before the stated cutoff and three weeks after the 16.0.0 the draw describes fluently. Model belief: "High that the segment cache is the mechanism and that `experimental.clientSegmentCache` is the flag ... The flag has taken values true, false, and 'client-only' across versions." Offered as the fix for prefetch request volume. Wrong: experimental: { clientSegmentCache: false } Correct: experimental: { prefetchInlining: true } Impact: Worse than a build error, because there is none. Unrecognised keys under `experimental` warn and are dropped — the build succeeds, the config reads as applied, and the prefetch burst the team was trying to fix continues unchanged. The draw's own hedge ("if it produces an unknown-option warning on your version, grep the release notes for the stabilised name") points at a rename that never happened: no config key in 16.3.4 contains the string 'segment'. Source: https://unpkg.com/next@16.0.2/dist/server/config-shared.d.ts (2025-11-12) — "clientSegmentCache?: boolean | 'client-only';" Source: https://unpkg.com/next@16.0.3/dist/server/config-shared.d.ts (2025-11-13) — "(no occurrence of clientSegmentCache in the file)" Source: https://nextjs.org/blog/next-16-2 (2026-03-18) — "The new experimental.prefetchInlining option bundles all segment data for a route into a single response, reducing the number of prefetch requests to one per link." --- F3 [S2 silently-wrong] Names Noto Sans as the ImageResponse default, having explicitly considered and rejected Geist API: ImageResponse default font Changed in next.js 16.2.0 (2026-03-18), kind: behavior-changed Chargeability: 16.2.0 published 2026-03-18, two months before the stated cutoff. Model belief: "Noto Sans — specifically Noto Sans Regular (weight 400, Latin subset) ... Confidence: moderate-to-high, not certain. I also have a vaguer, weaker recollection of discussion about switching next/og's default to Geist. If that switch shipped in a version I'm hazy on, the answer would be Geist Regular. I'd bet on Noto Sans." Impact: Every OG image generated without an explicit `fonts` option renders in Geist Sans. A team that lays out a card against Noto Sans metrics — and Noto Sans and Geist are not metrically compatible — gets different line breaks and overflow in the 1200x630 PNG than the model predicts. The interesting part is not the miss but the shape of it: the correct answer was present, weighed against the stale one, and lost. Source: https://unpkg.com/next@16.2.0/dist/compiled/@vercel/og/index.node.js (2026-03-18) — "var fontData = fs2.readFileSync( fileURLToPath(new URL("./Geist-Regular.ttf", import.meta.url)) );" Source: https://unpkg.com/next@16.1.7/dist/compiled/@vercel/og/index.node.js (2026-03-16) — "var fontData = fs2.readFileSync( fileURLToPath(new URL("./noto-sans-v27-latin-regular.ttf", import.meta.url)) );" --- F4 [S2 silently-wrong] Asserts Link has no transitionTypes prop and predicts a React unknown-attribute warning in both routers API: Link transitionTypes Changed in next.js 16.2.0 (2026-03-18), kind: added Chargeability: 16.2.0 published 2026-03-18, two months before the stated cutoff. Model belief: "I do not believe `transitionTypes` is a real prop on next/link in any stable Next.js I know of ... React DOM tries to set it as an attribute ... you end up with ... Yes, in both cases: Warning: React does not recognize the transitionTypes prop on a DOM element." Correct: About Impact: Three wrong answers in one: the prop exists and drives React's `addTransitionType` for that navigation; neither router forwards it to the ``, so no warning is logged in either; and a non-array value in development throws a Next prop-type error rather than producing an inert attribute. A team following this removes working view-transition code to silence a warning that was never there. Source: https://unpkg.com/next@16.3.4/dist/client/link.js (2026-08-31) — "legacyBehavior = false, transitionTypes, ...restProps } = props;" Source: https://nextjs.org/blog/next-16-2 (2026-03-18) — "transitionTypes on Pages Router links is silently ignored, so shared link components work across both routers." --- non-finding [context] error.tsx retry prop: DERIVABLE — task 1(b), the semantics of reset(). Both below-floor control subjects stated correctly and unprompted that reset() clears the boundary's error state and re-renders the already-downloaded RSC payload, that no request leaves the browser and that the server query does not re-run. This draw's identical answer is therefore not reported as knowledge of anything in the probe band. Pre-registered prediction P4 named this probe. --- non-finding [context] catchError (next/error): DERIVABLE — task 2(b), redirect() through a hand-rolled boundary. Both control subjects independently described the NEXT_REDIRECT digest being swallowed by a naive class boundary and named unstable_rethrow as the fix. The hazard predates the probe band, exactly as prediction P5 said. No pass on this probe is reported as knowledge. --- non-finding [miss] error.tsx retry prop: Task 1(a): asked to make a 'Try again' button actually re-run the server query, the draw hand-rolled useRouter().refresh() + reset() inside startTransition. That is exactly what the framework's own retry prop does — startTransition(() => { context.refresh(); reset() }) — which has been passed to every error.tsx since 16.2.0 (as unstable_retry, stable as retry in 16.3.0). Not charged: the hand-rolled code works. The Index's four-level severity scale has no slot for correct code that a framework API now supersedes, and inventing one to book this would be worse than leaving it uncharged. Recorded so the count is honest. --- non-finding [miss] catchError (next/error): Task 2(a): offered parallel routes with a per-slot error.tsx, then a hand-rolled React class boundary with unstable_rethrow. Both work; neither is catchError from next/error, which has existed since 16.2.0 and is designed for exactly this. Same scale problem as task 1(a) — recorded, not charged. --- non-finding [correct] images.maximumDiskCacheSize: Task 5(a), the half the battery actually cared about: 'recent Next.js caps the optimized-image cache as a fraction of free space on the volume holding .next/cache, measured once when the server process starts — not a fixed byte count, and not re-measured as the disk fills', with LRU eviction. That is the rule, the source (free rather than total space) and the timing, all correct, for a default that shipped in a patch release two months before this subject's cutoff and which no control subject described. The draw declined to guess the fraction and then guessed 10%; the real figure is 50%. Pre-registered prediction P2 said this probe would come back 'unbounded' or misdirected to minimumCacheTTL. P2 is falsified. --- non-finding [context] images.maximumDiskCacheSize: The twin disagrees on the number. Asked the identical question from the same stored prompt, v2-b gave 50% — the correct figure — while this draw gave 10%. Both gave the correct rule and timing. The arm that charges is the one that missed the number, which is the third battery running in which the -b twin holds the better answer on some probe; see the undercount note in HARNESS.md. --- non-finding [context] images.maximumResponseBody: VOID PROBE — task 6. The task said 'default image configuration' while handing over a remote src, and three of four draws reasonably read that as remotePatterns being unset and answered 400-host-not-allowed. The probe cannot distinguish a subject that knows the 50 MB body limit from one that stopped at the allowlist, and it is scored for nobody. This draw did volunteer, unprompted, that 'there is an upstream size limit above which it refuses to buffer and optimize the source' without naming a figure. The battery's wording is the fault, not the answer. --- non-finding [imprecision] images.qualities: Task 7, the internal control. Answered that q=90 with no images.qualities is rejected with a 400, then hedged explicitly to the correct alternative — 'If it's the latter, the answer to (a) is 75 — your quality={90} is ignored' — which the code-vs-claim rule records as an imprecision rather than a finding. The attribution half was correct and unhedged: 16.0.0, 2025-10-21. The control did its job: this subject's attribution answers are readable, so its failures elsewhere are failures of knowledge and not of dating. --- non-finding [context]: Task 9, attribution. The OG font default was placed at 13.3.0 (April 2023) and the image disk-cache rule at 16.0.0, with 16.1 as the alternative. The true answers are 16.2.0 and 16.1.7 — a minor and a patch published two days apart. No draw of four reached either, and none reached any patch. But this battery cannot claim better-auth's result: there, subjects held a capability and misplaced it; here three of the four attribution targets were behaviours the subject did not hold, so the misdating follows from the gap rather than measuring attribution independently. The one attribution question asked about a behaviour this subject does hold — task 7, the qualities default — was answered correctly, and that change shipped in a major. Two libraries now point the same way: attribution survives majors and fails on patches. --- non-finding [context]: The boundary reproduced across two different batteries. next.js/v1 (2026-08-31) measured this subject's attribution boundary at 16.0.0 / 16.1.0 with a completely different prompt; v2 lands on the same pair two days later, in its own words ('partial and unreliable knowledge of 16.1'). Every previously published boundary agreement came from replicates of one prompt. Not re-charged: the S4 recency finding is v1's F4 and stands there. ============================================================================== RUN next.js--claude-opus-5--v2-b--2026-09-02 What Claude Opus 5 gets right about next.js — battery v2-b, tested 2026-09-02 URL: https://stalepriors.com/runs/next.js--claude-opus-5--v2-b--2026-09-02 JSON: https://stalepriors.com/data/next.js/opus-5-v2-b.json Library: next.js 16.3.4 (npm), verified 2026-09-02 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 16.0.0 (2025-10-22), lag ~7 months Oldest release it could not place: 16.1.0 (2025-12-18) Battery: next.js/v2-b, 9 tasks, tool uses during test: 0 Tested: 2026-09-02 Summary: The second, blind draw of the duplicated test arm. It charges nothing by rule. It agreed with its twin on every probe outcome — same four failures, same boundary, same hand-rolled answers to both capability probes — with one exception that matters: asked what fraction of the disk the optimized-image cache may use, this draw said 50%, which is right, where the charging twin said 10%. Two draws of one prompt, one number apart, and the wrong one is the one that counts. That is the third battery in which the -b twin holds the better answer on some probe. --- non-finding [correct] images.maximumDiskCacheSize: Task 5(a), and the reason this run exists: 'the cache is bounded to a fraction of the free space on the volume holding .next/cache, measured once, lazily, when the image optimizer initialises ... I believe the fraction is 50%'. The rule, the source, the timing and the figure are all correct, for a default that shipped in the patch release 16.1.7. Its twin gave the same rule and guessed 10%. The battery's pre-registered prediction P2 — that this probe would come back 'unbounded' or misdirected to minimumCacheTTL — is falsified twice over. --- non-finding [miss] images.maximumCacheSize: Task 5(b)/(c): wrote images.maximumCacheSize, the same invented key as its twin, having correctly stated that 'an unknown images.* key is a hard config validation error at boot'. Charged as F1 on v2-a; recorded here as a chargeable miss so the pair reads honestly. --- non-finding [miss] experimental.clientSegmentCache: Task 8: experimental.clientSegmentCache: false, described as tri-state and default-on in the 16 line. The key was removed in 16.0.3, one day after the 16.0.2 that still had it, and unrecognised experimental keys warn rather than fail — so the recommendation is inert. Charged as F2 on v2-a. --- non-finding [miss] ImageResponse default font: Task 3: 'Noto Sans — specifically Noto Sans Regular, the Latin subset (the file shipped/fetched as noto-sans-v27-latin-regular.ttf)', at high confidence. That filename is exactly right for 16.1.7 and exactly wrong for 16.2.0, which replaced it with Geist-Regular.ttf two days later. Where the twin weighed a Geist recollection and rejected it, this draw did not surface one at all. Charged as F3 on v2-a. Both draws of the test arm and both control draws named Noto Sans: pre-registered prediction P1 holds, four for four. --- non-finding [miss] Link transitionTypes: Task 4: put 65% on transitionTypes not being a real prop, predicted the array would be stringified onto the anchor and that React would warn in both routers. It then described the real behaviour accurately as its 35% branch — 'it would tag the client-side navigation's view transition with the type slide ... in that case (c) flips to no warning in either'. The correct answer was reachable and was priced at a third. Charged as F4 on v2-a. --- non-finding [miss] error.tsx retry prop: Tasks 1(a) and 2(a): hand-rolled router.refresh() + reset() in a transition, and a hand-rolled class boundary with unstable_rethrow. Identical to the twin, and identically unchargeable — the code works, and the scale has no level for 'superseded by a first-class API'. --- non-finding [context]: DERIVABLE, both: task 1(b) reset() semantics and task 2(b) redirect() through a naive boundary. Both were answered correctly by both below-floor control subjects, so neither pass is reported as knowledge of anything in the probe band. --- non-finding [context] images.maximumResponseBody: VOID PROBE — task 6, whose wording let 'default image configuration' be read as remotePatterns being unset. This draw hedged across both readings and put ~55% on an upstream size guard existing without naming 50 MB. Scored for nobody. --- non-finding [imprecision] images.qualities: Task 7, the internal control: same shape as the twin — 400 asserted, clamping-to-75 named as the alternative at moderate confidence, and the release given correctly as 16.0 on 21 October 2025, with images.qualities correctly dated to 15.3 as an unrestricted opt-in. Attribution is readable for this subject. --- non-finding [context]: Task 9: the OG font default placed at 13.0.0 (October 2022) and the disk-cache rule at 16.0.0. Its twin said 13.3.0 and 16.0.0. Neither reached 16.2.0 or 16.1.7; no draw of four reached any patch release. ============================================================================== RUN next.js--claude-opus-5--v3-e--2026-09-02 What Claude Opus 5 gets wrong about next.js — battery v3-e, tested 2026-09-02 URL: https://stalepriors.com/runs/next.js--claude-opus-5--v3-e--2026-09-02 JSON: https://stalepriors.com/data/next.js/opus-5-v3-e.json Library: next.js 16.3.4 (npm), verified 2026-09-02 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 16.0.0 (2025-10-22), lag ~7 months Oldest release it could not place: 16.1.0 (2025-12-18) Battery: next.js/v3-e, 5 tasks, tool uses during test: 0 Tested: 2026-09-02 Summary: The arm that was supposed to be a repeat and was not. Asked the same thing v2 asked, in tighter wording, this subject reversed itself: where all four v2 draws recommended experimental.clientSegmentCache, this draw answered "No" in one word, stated the key is gone in 16, and predicted the exact unrecognised-key warning the shipped config emitter produces. That answer is right, and it does not clear the subject - having correctly buried the dead key it then denied that any lever exists, when experimental.prefetchInlining (16.2.0) does precisely what the task asked for. This is the right-about-the-removal, wrong-about-the-replacement category, charged as F1, and it is the second library where the Index has caught it. Task 3 is charged as F2: it leaned to a cap existing at 60/40, put it in the "low tens of megabytes", denied a config key exists and declined to name one - the truth is 50 MB, a 413, and images.maximumResponseBody. On the internal control it named proxy.ts and dated the rename to 16.0.0 correctly, so its attribution answers are readable; and having established that, it placed the clientSegmentCache removal in the 16.0.0 major when it shipped in the 16.0.3 patch. That is the patch-granularity result for the third time, now from a subject that demonstrably holds the change. --- F1 [S2 silently-wrong] Denies any lever on prefetch request count exists, three releases after experimental.prefetchInlining shipped one API: experimental.prefetchInlining Changed in next.js 16.2.0 (2026-03-18), kind: added Chargeability: Published 2026-03-18, two months before this draw's stated cutoff of May 2026, which it accepted. Not probed as a finding by v2: v2's task 8 charged the dead key this draw correctly rejects, and never reached the question of what replaced it. Model belief: Task 1(a), one word: "No". Task 1(b): "There is no config key, no prop, and no runtime API that says 'give me one whole-tree prefetch response per link' any more." It then offered five alternatives - prefetch={false}, a hand-rolled intent-delay wrapper, flattening the route tree, an argument that the team's premise is wrong, and pinning to 15.x - and closed with "high (~85%) that there is no supported off switch on 16". Wrong: Go Correct: experimental: { prefetchInlining: true } Impact: The team asked for one prefetch request per link. There is a config flag that does exactly that, and they are told it does not exist, then handed a client component that delays prefetching, advice to restructure their layouts, and a suggestion to pin an old major. Every one of those is real work; the supported answer is one line. This is the failure mode the capability probe exists to catch - the removal is known, the replacement is not, so the subject argues from the gap rather than reporting it. Scope note: prefetchInlining is still flagged experimental at 16.3.4. The finding is the denial that any lever exists, not a claim that a stable API was missed. Source: https://nextjs.org/blog/next-16-2 (2026-03-18) — "The new experimental.prefetchInlining option bundles all segment data for a route into a single response, reducing the number of prefetch requests to one per link." Source: https://unpkg.com/next@16.3.4/dist/server/config-shared.d.ts (2026-08-31) — "prefetchInlining?: boolean | {" --- F2 [S2 silently-wrong] Puts the image body cap in the low tens of megabytes and denies a config key lowers it API: images.maximumResponseBody Changed in next.js 16.1.5 (2026-01-26), kind: added Chargeability: Published 2026-01-26, four months before this draw's stated cutoff. The probe was rewritten for this battery because v2's version of it was void - it said "default image configuration" while handing over a remote src, and three of four draws answered about the remotePatterns allowlist instead. This version states the whole config file and stipulates that remotePatterns matches. Model belief: "My belief is that the cap is a fixed constant in the low tens of megabytes - if forced to a single figure I would say on the order of 10 MB" and "I do not believe there is a documented, stable images.* key that lowers it. I specifically decline to invent a key name here." It leaned to a 400 status at 60/40. The real answer is 50 MB, a 413, and images.maximumResponseBody. Wrong: images: { loader: 'custom', loaderFile: './image-loader.ts' } // plus an /api/img route doing a HEAD content-length check Correct: images: { maximumResponseBody: 5_000_000 } Impact: The one-line config the task asked for is replaced by a custom loader, a route handler, a HEAD request per image and an acknowledged race ("a hostile origin can lie"). The refusal to invent a key name is the right instinct and is recorded as such; the charge is on the denial that the key exists, which sends a reader who could have set one option down a path that reimplements the optimizer's own check less well. Scope note: Verified in the published packages, not executed against a live 80 MB fetch. The 413 and the 50 MB constant are read from the shipped optimizer and the shipped default config. Source: https://unpkg.com/next@16.3.4/dist/shared/lib/image-config.js (2026-08-31) — "maximumResponseBody: 50000000," Source: https://unpkg.com/next@16.3.4/dist/server/image-optimizer.js (2026-08-31) — "if (totalSize > maximumResponseBody) {" Source: https://unpkg.com/next@16.1.5/dist/shared/lib/image-config.js (2026-01-26) — "maximumResponseBody: 50000000," --- non-finding [correct] experimental.clientSegmentCache: Tasks 1(a) and 2, and the reason this battery is not a repeat of v2. Every draw of v2 recommended experimental.clientSegmentCache; this draw refused it - "The thing that used to control it (experimental.clientSegmentCache) is gone" - answered task 2(a) with "It prints a warning and continues", reproduced the unrecognised-key warning almost verbatim, and answered 2(b) "Nothing" with the correct mechanism: "in 16 the segment cache is the router, there is no other code path, and no code reads that key any more." Correct on every part. Whether the reversal is knowledge or wording is the open question below. --- non-finding [correct] middleware.ts / proxy.ts: Task 4, the internal control, passed cleanly: proxy.ts exporting proxy, with middleware.ts named as the older spelling and correctly described as still working with a deprecation warning. Task 5(c) then dated the rename to 16.0.0, 21 October 2025, correct to the release. This subject holds the change and can place it in the major, which is what makes its failure on 5(a) readable as an attribution result rather than a knowledge gap. --- non-finding [miss] experimental.clientSegmentCache: Task 5(a), the patch-granularity probe. Dated the clientSegmentCache removal to "Next.js 16.0.0, approximately 21 October 2025" and rated it moderate confidence, explicitly weighing and rejecting the alternative that it survived into 16.0 and was deleted later. It shipped in 16.0.3, three weeks after 16.0.0. The subject holds the change - it described the key's removal correctly in tasks 1 and 2 - and still cannot reach the patch. Not charged: the version-attribution S4 rule charges recency against the subject's own boundary, and this is the battery's designed attribution reading rather than a separate belief failure. --- non-finding [correct] images.maximumResponseBody: Task 3(c), answered correctly and for the correct reason while the rest of task 3 was wrong: "If a cap applies it is a fixed byte constant compiled into the optimizer, not a function of os.freemem() ... a 256 GB machine gets the same 400 as a 512 MB one." Its blind twin v3-f answered the same sub-question the other way, calling machine-dependence "the diagnostic signature of no hard limit". The discriminator worked: it separated the fixed-policy world from the resource world, and the twins landed on opposite sides of it. ============================================================================== RUN next.js--claude-opus-5--v3-f--2026-09-02 What Claude Opus 5 gets right about next.js — battery v3-f, tested 2026-09-02 URL: https://stalepriors.com/runs/next.js--claude-opus-5--v3-f--2026-09-02 JSON: https://stalepriors.com/data/next.js/opus-5-v3-f.json Library: next.js 16.3.4 (npm), verified 2026-09-02 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 16.0.0 (2025-10-22), lag ~7 months Oldest release it could not place: 16.1.0 (2025-12-18) Battery: next.js/v3-f, 5 tasks, tool uses during test: 0 Tested: 2026-09-02 Summary: The blind twin, and it agrees with v3-e everywhere the twin pair matters and splits from it where the battery was built to look. It gave the same one-word "No" on task 1, the same correct warning-and-continues on task 2(a), and a better answer than its twin on 2(b) - spelling out that the config is a no-op on 16 because the key is unrecognised and a no-op on 15.x because false was the default, then naming the trap: "on 16 it looks like an opt-out and silently is not one". It repeated the twin's denial that any prefetch lever exists, which is the F1 failure carried on the charging arm. On task 3 the twins diverge cleanly: v3-e leaned to a cap and correctly said available memory is irrelevant to it; this draw leaned to success at 200, no cap, and answered the RAM question "Yes - and that is exactly the tell", reading machine-dependence as evidence against a fixed limit. It is the wrong side of a discriminator that worked. Both draws then dated the clientSegmentCache removal to the 16.0.0 major, three weeks early, and the proxy.ts rename to 16.0.0, correct. --- non-finding [miss] experimental.prefetchInlining: Task 1: the same denial as the charging twin - "my understanding is that the segment cache is on by default and there is no supported config key to revert to whole-tree prefetching" - followed by hand-rolled intent prefetching, layout flattening, an argument against the team's premise, and pinning to 15.x. experimental.prefetchInlining (16.2.0, two months inside this subject's stated cutoff) does exactly what was asked. Charged on the twin as F1; this is the -b-role draw of the duplicated arm and charges nothing. --- non-finding [miss] images.maximumResponseBody: Task 3: answered that the request ends 200 with a normally optimized WebP and that no configurable byte cap exists - "There is no images.maxUpstreamSize / images.maximumFileSize that I can attest to" - listing the image options it is confident of and correctly noting maximumRedirects among them while missing maximumResponseBody, which sits beside it in the same shipped default config. The real answer is a 50 MB cap enforced while streaming, a 413, and images.maximumResponseBody. Charged on the twin as F2; this arm charges nothing. --- non-finding [correct] experimental.clientSegmentCache: Task 2, the sharpest answer any draw of six gave on the dead key. "On Next.js 16: the key is unrecognised, so it is discarded during config validation and never reaches the router. Writing false does not turn the segment cache off - an unknown key is inert, not an override." It separated that from the 15.x reading, where the key is recognised and false is the default, and named the consequence: "a developer reading this config would reasonably conclude the app is on whole-tree prefetching when it is not." Both Opus draws reversed v2's four-of-four failure on this surface under the tighter wording. --- non-finding [imprecision] proxy.ts: Task 4(a): named proxy.ts correctly but specified a default export where the convention is a named export proxy - then hedged it and named the correct alternative in the same breath: "Medium on the default-vs-named export detail (I am fairly sure it is a default export; if your build complains, a named export function proxy is the thing to try)." Recorded rather than charged under the code-vs-claim rule, which treats a hedged claim that names the correct fix as an imprecision. --- non-finding [miss] images.maximumResponseBody: Task 3(c), the discriminator, answered on the opposite side from its twin: "Would more RAM and disk change the answer? Yes - and that is exactly the tell ... That non-determinism is the diagnostic signature of no hard limit." The reasoning is sound and the premise is false: the limit is a fixed constant in imageConfigDefault, so the outcome is machine-independent. Not separately chargeable - it is the same belief as the task 3 miss above, stated as an inference. --- non-finding [miss] experimental.clientSegmentCache: Task 5(a): "Next.js 16.0.0, approximately 21 October 2025", labelled a guess, with the alternative considered and rejected. Actual: 16.0.3, 2025-11-13. Both Opus draws collapsed the patch onto the major while both placed the 16.0.0 rename correctly in task 5(c). Two subjects, three batteries, and no draw has yet attributed a change to a patch release. ============================================================================== RUN next.js--claude-opus-5--v4-e--2026-09-05 What Claude Opus 5 gets right about next.js — battery v4-e, tested 2026-09-05 URL: https://stalepriors.com/runs/next.js--claude-opus-5--v4-e--2026-09-05 JSON: https://stalepriors.com/data/next.js/opus-5-v4-e.json Library: next.js 16.3.4 (npm), verified 2026-09-05 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 16.0.0 (2025-10-22), lag ~7 months Oldest release it could not place: 16.1.0 (2025-12-18) Battery: next.js/v4-e, 0 tasks, tool uses during test: 0 Tested: 2026-09-05 Summary: The bridge arm in the suppressing order: same subject as `prisma/v3`, new library, boundary question FIRST. On prisma that order cost this subject thirteen releases of demonstrated knowledge. Here it costs nothing measurable - D = 16.0.0 and S = 16.0.0/16.1.0, which is exactly what its `-cs` twin `v4-f` produced and exactly what five prior next.js runs of this subject published. Poison and ceiling controls clean. --- non-finding [correct] next 16.0.0: LADDER GRADE - CORRECT, under a wrapper hedge that HARNESS.md clause 2 says does not void it. Opened "moderate confidence on the headline set, lower on individual items" and then named Turbopack as the default bundler with a webpack opt-out, Cache Components with the `use cache` directive and a `cacheComponents` config flag as successor to `ppr`/`dynamicIO`, `updateTag`/`refresh` alongside `revalidateTag`, `middleware.ts` to `proxy.ts`, React 19.2, the AMP and `next lint` removals and the Node 20.9 floor. Every one is in the release post. --- non-finding [context] next 16.1.0: LADDER GRADE - ABSTAIN, and the cleanest instance in the battery of an unbound fragment. "I have fragments I loosely associate with the post-16.0 window - a Next.js devtools MCP server, Turbopack filesystem caching, and an experimental deployment adapters API - but I cannot tell you which release any of those shipped in, or whether they shipped at all." Two of the three are real and one of them (file-system caching for `next dev`) is 16.1's headline, but the arm explicitly refuses to bind any of them to a release, which is exactly the shape clause 2 excludes from D. --- non-finding [correct] next 15.3.0: LADDER GRADE - CORRECT. Turbopack production builds alpha, experimental `next-rspack`, `instrumentation-client.ts` running before app code, `onNavigate` on Link, `useLinkStatus`, TypeScript language-server performance. --- non-finding [correct] next 15.1.0: LADDER GRADE - CORRECT. React 19 stable, `after()` graduated from `unstable_after`, `forbidden()`/`unauthorized()` behind `experimental.authInterrupts` with both file conventions, and React owner stacks in the dev overlay. --- non-finding [correct] next 15.5.0: LADDER GRADE - CORRECT. Turbopack production builds beta, Node.js middleware runtime stable, `typedRoutes` stable, `next typegen`, the `PageProps`/`LayoutProps`/`RouteContext` helpers, and the 16 deprecation warnings. --- non-finding [correct] next 15.7.0: POISON RUNG. 15.7.0 has never been published in any form - zero versions of any kind match 15.7.* on the npm registry, checked this session. "I have no record of a 15.7.0 at all... I would not assert 15.7.0 was never released - it could be a post-16 maintenance release outside my knowledge - but I can attach nothing to it." It declines without inventing content, which is what the control requires. It also reports "a hazy sense of a 15.6 shipping as a bridge release": no stable 15.6 was ever published, only 62 canaries, so that sense is of a version that exists in prereleases alone. P4 holds on this arm, so its CORRECT grades stand. --- non-finding [correct] next 16.3.0: CEILING RUNG. 16.3.0 (2026-08-03) is above this subject's stated cutoff. "Cannot describe. If it exists, it would land in roughly the first half of 2026, at or past the edge of my usable knowledge." P5 holds on this arm. --- non-finding [context] self-placement: Question (a) and question (c) AGREE on this subject: (a) names 16.0.0 as the most recent describable release and (c) names 16.1.0 as the first release known only as a version number. The one-release divergence seen on all four Claude Fable 5.1 arms is absent here, so it is a property of that subject rather than of the question pair. ============================================================================== RUN next.js--claude-opus-5--v4-f--2026-09-05 What Claude Opus 5 gets right about next.js — battery v4-f, tested 2026-09-05 URL: https://stalepriors.com/runs/next.js--claude-opus-5--v4-f--2026-09-05 JSON: https://stalepriors.com/data/next.js/opus-5-v4-f.json Library: next.js 16.3.4 (npm), verified 2026-09-05 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 16.0.0 (2025-10-22), lag ~7 months Oldest release it could not place: 16.1.0 (2025-12-18) Battery: next.js/v4-f, 0 tasks, tool uses during test: 0 Tested: 2026-09-05 Summary: The bridge arm in the favourable order: ladder first, boundary question last, the position every published battery uses. D = 16.0.0, S = 16.0.0/16.1.0 - identical to `v4-e`, which was asked the same questions in the opposite order. The order effect that moved this subject thirteen releases on prisma moves it zero releases here. Poison and ceiling controls clean, and its ceiling answer names its own cadence inference and then declines to trade on it. --- non-finding [correct] next 16.0.0: LADDER GRADE - CORRECT. Turbopack default for dev and build, Cache Components graduating `dynamicIO` into a `cacheComponents` option with `use cache`, `cacheLife`/`cacheTag`, `updateTag`/`refresh` beside `revalidateTag`, `proxy.ts` replacing `middleware.ts`, React 19.2 with Activity, the AMP and `next lint` removals, the Node 20.9 floor and the adapters API in alpha. Self-rated "high" on most of that list and "medium" on `proxy.ts` and the adapters API - both of which are correct. --- non-finding [context] next 16.1.0: LADDER GRADE - ABSTAIN, unambiguous. "Cannot describe. I believe a 16.1 was released (the 16.x line clearly continued past the initial release), but I cannot name a single feature, config key, or API from it with any confidence. I'd rather say nothing than guess." Existence without content, which is the definition the Index's boundary interval is built on. --- non-finding [correct] next 15.3.0: LADDER GRADE - CORRECT. `next build --turbopack` alpha, `instrumentation-client.ts`, `onNavigate` and `useLinkStatus`, experimental `next-rspack`, TypeScript language-server performance. --- non-finding [correct] next 15.1.0: LADDER GRADE - CORRECT. React 19 stable, `after()` stable with the correct import path, `forbidden()`/`unauthorized()` behind `experimental.authInterrupts` with `forbidden.tsx`/`unauthorized.tsx`, and the source-map and stack-trace work. --- non-finding [correct] next 15.5.0: LADDER GRADE - CORRECT. Turbopack production builds beta, Node.js middleware stable, `typedRoutes` out of experimental with route export validation, `next lint` deprecated with removal announced for 16, and the AMP deprecation warning. --- non-finding [correct] next 15.7.0: POISON RUNG. 15.7.0 has never been published in any form - zero versions of any kind match 15.7.* on the npm registry, checked this session. "Cannot describe. I have no content for this version, and I am not confident it was ever released - my knowledge of the 15.x line effectively ends at 15.5 with a possible 15.6 that I also cannot describe." P4 holds on this arm, so its CORRECT grades stand. --- non-finding [correct] next 16.3.0: CEILING RUNG. 16.3.0 (2026-08-03) is above this subject's stated cutoff. "Cannot describe... I would guess from release cadence that a 16.3 plausibly exists by mid-2026, but that is inference from the pattern, not knowledge, so I am not going to characterize it." Naming the inference and then refusing to trade on it is the behaviour the ceiling rung exists to check for. P5 holds on this arm. --- non-finding [context] self-placement: "First release I know only as a version number, with no content attached: 16.1.0." Identical to its `-sc` counterpart `v4-e` and to the five prior next.js runs of this subject. The boundary this Index publishes for Claude Opus 5 on next.js is now seven runs deep and has never moved. ============================================================================== RUN next.js--claude-opus-5--v1--2026-08-31 What Claude Opus 5 gets wrong about next.js — battery v1, tested 2026-08-31 URL: https://stalepriors.com/runs/next.js--claude-opus-5--v1--2026-08-31 JSON: https://stalepriors.com/data/next.js/opus-5.json Library: next.js 16.3.3 (npm), verified 2026-08-31 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 16.0.0 (2025-10-22), lag ~7 months Oldest release it could not place: 16.1.0 (2025-12-18) Battery: next.js/v1, 10 tasks, tool uses during test: 0 Tested: 2026-08-31 Summary: Knows Next.js 16 exists and gets most of it right — proxy.ts, next lint removal, Turbopack as the default bundler, the image defaults — yet still ships two pieces of code that do not run on it, and misdates the proxy rename by a full major version. Four findings, all chargeable, two S1. The result that matters is comparative: this subject has the latest stated cutoff in the dataset (2026-05) and is still beaten on this library by Fable 5, whose stated cutoff is four months earlier. Disclosed self-test — the operator model is the subject. --- F1 [S1 breaks-build] Ships experimental.ppr and experimental_ppr as the primary answer, both removed API: experimental.ppr / export const experimental_ppr Changed in next.js 16.0.0 (2025-10-22), kind: removed Model belief: Gave the removed configuration as the working answer across three files, then added: "in Next.js 16 this machinery was folded into the cacheComponents: true flag (renamed from the earlier dynamicIO) ... The experimental.ppr + experimental_ppr form above is what I can describe confidently for 15.x; check whether 16 still accepts it in your version." Wrong: // next.config.ts const nextConfig: NextConfig = { experimental: { ppr: 'incremental' } }; // app/marketing/page.tsx export const experimental_ppr = true; Correct: // next.config.ts const nextConfig: NextConfig = { cacheComponents: true }; // app/marketing/page.tsx — no route-segment export needed Impact: It does not accept it. Both were removed in 16.0.0. The model names the correct replacement in prose and then hands over the removed one as the code, which is the failure mode this battery is built to catch: knowing a fact and not applying it. Source: https://nextjs.org/blog/next-16 (2025-10-21) — "`experimental.ppr` flag | PPR flag removed; evolving into Cache Components programming model" Source: https://nextjs.org/blog/next-16 (2025-10-21) — "`export const experimental_ppr` | Route-level PPR export removed; evolving into Cache Components programming model" --- F2 [S1 breaks-build] Single-argument revalidateTag in a generated Server Action API: revalidateTag Changed in next.js 16.0.0 (2025-10-22), kind: behavior-changed Model belief: Wrote `revalidateTag('products')` in the action. At question (c): "revalidateTag arguments: one — the tag string ... The two-argument shape people are thinking of is revalidatePath." It then added: "I have a faint recollection of an optional second cache-profile argument on revalidateTag appearing in the Next.js 16 line ... I am not confident in that and won't assert it." Wrong: 'use server' import { revalidateTag } from 'next/cache' export async function updateProduct(id: string, formData: FormData) { await db.product.update({ where: { id }, data: { ... } }) revalidateTag('products') redirect(`/products/${id}`) } Correct: 'use server' import { revalidateTag, updateTag } from 'next/cache' export async function updateProduct(id: string, formData: FormData) { await db.product.update({ where: { id }, data: { ... } }) revalidateTag('products', 'max') // stale-while-revalidate // or updateTag('products') for read-your-writes inside the action redirect(`/products/${id}`) } Impact: TypeScript project: the single-argument form is documented as producing a TypeScript error, and `next build` type-checks by default. The model had the right answer available — it names `updateTag` and `refresh` correctly two paragraphs earlier — and specifically declined to trust it. Source: https://nextjs.org/docs/app/guides/upgrading/version-16 — "revalidateTag now requires a second argument specifying a cacheLife profile. The single-argument form is deprecated and will produce a TypeScript error." --- F3 [S4 wrong-metadata] Dates the middleware-to-proxy rename to 15.5; it shipped in 16.0.0 API: proxy.ts Changed in next.js 16.0.0 (2025-10-22), kind: version-fact Model belief: "through Next.js 15.4 this file was middleware.ts. In 15.5 the proxy.ts name was introduced and middleware.ts began being deprecated; in Next.js 16 proxy.ts is the documented convention." It repeats the boundary twice more: "rename to middleware.ts if you're on Next.js ≤ 15.4" and "On 15.4 and earlier, middleware.ts remains correct." Impact: A team on 15.5 following this advice renames to proxy.ts, where nothing picks the file up, and the route gate silently stops running — the worst class of auth bug, because the app still serves. The rename landed in 16.0.0 and nowhere earlier. Note the same answer misses that the proxy runtime is nodejs and cannot be configured, which matters for anyone on an edge deployment. Source: https://nextjs.org/blog/next-16 (2025-10-21) — "`middleware.ts` filename | Rename to `proxy.ts` to clarify network boundary and routing focus" Source: https://nextjs.org/docs/app/guides/upgrading/version-16 — "The edge runtime is NOT supported in proxy. The proxy runtime is nodejs, and it cannot be configured." --- F4 [S4 wrong-metadata] Version knowledge stops at 16.0.0, seven months before its stated cutoff Changed in next.js 16.1.0 (2025-12-18), kind: version-fact Chargeability: Anchored to 16.1.0 (2025-12-18), the first release the model cannot describe — not to the current release. 16.1.0 shipped five months before this model's stated 2026-05 cutoff. Model belief: "Versions I'm aware of but can't describe: I believe there have been 16.x point releases after that — 16.1 around December 2025, and probably more in early 2026 — but I cannot tell you what's in them, and I'd be fabricating if I named a specific 'latest' version number." Impact: Two further minors (16.2.0, 2026-03-18) and a long tail of patches shipped inside this model's training window and are absent from it. The self-diagnosis is exact — "knowledge density drops sharply in the last few months before any cutoff" — and the run measures that drop at seven months on this library. The candour is correct behaviour and is noted in its favour; the gap is still the finding. Source: https://api.github.com/repos/vercel/next.js/releases (2025-12-18) — ""tag_name": "v16.1.0", "published_at": "2025-12-18T18:49:40Z"" --- non-finding [correct] proxy.ts: Task 1: named proxy.ts and the proxy export, correctly identifying the file rename that Sonnet 5 misses entirely. Only the version attribution is wrong (F3). --- non-finding [correct] next lint / next build: Task 6: "next lint was deprecated in Next.js 15.5 and removed in Next.js 16" and generated "lint": "eslint". Also correct that next build uses Turbopack by default in 16 with `--webpack` as the opt-out. --- non-finding [correct] cacheLife / cacheTag: Task 5: offered both `unstable_cache` and the `'use cache'` form with unprefixed `cacheLife` / `cacheTag` — correct for 16.0.0, which stabilized both. --- non-finding [correct]: Task 3: async `params` in the OG image route. Task 7: remotePatterns. Task 8: knew `images.qualities` defaults to `[75]` and that 90 is not served. Task 9: listed `@panel/default.tsx`. --- non-finding [imprecision] images.minimumCacheTTL: Task 10: gave the correct new default (4 hours / 14400s) but attributed it to "a Next.js 15.4-era release" rather than 16.0.0. The value is right and the advice — set it explicitly — is sound, so this is recorded rather than charged; the analogous misdating on the proxy rename is charged as F3 because acting on it silently breaks a route gate. --- non-finding [imprecision] images.qualities: Task 8: uncertain whether an unlisted quality is coerced or rejected with a 400 — "I've seen both described." The upgrade guide says coerced to the nearest allowed value. Explicitly hedged and paired with the correct fix, so not charged under the code-vs-claim rule. --- non-finding [imprecision] images.domains: Task 7: says `images.domains` is "deprecated (and I believe removed in 16)". It is deprecated in 16.0.0, not removed. No consequence — the model used remotePatterns. --- non-finding [context]: Disclosed self-test: the operator model for this studio is Claude Opus 5, the same model under test here. The run is weaker evidence than the other two for that reason. It is included because excluding the operator's own model from a dataset about model staleness would be the more dishonest choice, and because the finding it produces is unflattering. ============================================================================== RUN next.js--claude-sonnet-5--v2-c--2026-09-02 What Claude Sonnet 5 gets right about next.js — battery v2-c, tested 2026-09-02 URL: https://stalepriors.com/runs/next.js--claude-sonnet-5--v2-c--2026-09-02 JSON: https://stalepriors.com/data/next.js/sonnet-5-v2-c.json Library: next.js 16.3.4 (npm), verified 2026-09-02 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01 Version attribution stops at: 15.0.0 (2024-10-21), lag ~15 months Oldest release it could not place: 15.1.0 (2024-12-10) Battery: next.js/v2-c, 9 tasks, tool uses during test: 0 Tested: 2026-09-02 Summary: A below-floor control arm. This subject's next.js attribution stops at 15.0.0 — fifteen months before its own stated cutoff and sixteen releases below the probe band — so it charges nothing and exists to mark which probes are answerable without knowing the band. It struck two: the semantics of reset() and the redirect()-through-a-boundary hazard, both answered correctly and in detail. It also came within one word of striking a third, guessing 50% for the image disk-cache budget while getting the rule and the timing wrong. Its most useful contribution is negative: on the battery's internal control it dated the images.qualities default change to 15.3, not 16.0, so this subject's attribution answers are unreadable and no dating question in this battery can be scored against it. --- non-finding [context] error.tsx retry prop: DERIVABLE — task 1(b). Answered correctly and unprompted from fifteen months below the band: 'reset() is purely a React error-boundary state reset ... does not talk to the network, does not invalidate the Next.js Router Cache, and does not re-invoke the Server Component ... the database query does not run again'. No pass on this probe by any arm is reported as knowledge. This is pre-registered prediction P4, confirmed. --- non-finding [context] catchError (next/error): DERIVABLE — task 2(b). Described the NEXT_REDIRECT digest being swallowed by a hand-rolled boundary and named unstable_rethrow as the fix, from below the floor. Prediction P5, confirmed. --- non-finding [context] images.maximumDiskCacheSize: PARTIALLY DERIVABLE — task 5(a). Reached the figure without the rule: '50% of the volume's total disk space, checked at the point a new optimized variant is about to be written (using a disk-space check, not computed once at server boot)'. The true rule is 50% of *free* space computed *once at startup*, and the task explicitly asked when it is computed, so this is not a pass. But the number is evidently reachable from below the floor, which downgrades the number half of the test-arm's answer. The rule-and-timing half is not derivable: this draw asserted the opposite of it, and the other control denied any such cap exists. --- non-finding [miss] ImageResponse default font: Task 3: Noto Sans, at ~75% confidence. Four draws of four named Noto Sans. Not chargeable — 16.2.0 is fourteen months above this subject's attribution boundary and two months above its stated cutoff. --- non-finding [miss] Link transitionTypes: Task 4: asserted transitionTypes is not a recognised prop, predicted the React unknown-attribute warning in both routers. Same wrong answer as every other draw. Not chargeable — above this subject's cutoff. --- non-finding [miss] experimental.clientSegmentCache: Task 8: experimental.clientSegmentCache: false, at ~40% confidence on the key name. The key was removed in 16.0.3 on 2025-11-13 — two months INSIDE this subject's stated cutoff, and the failure is therefore chargeable in principle. It is not charged, because a control arm charges nothing. Queued: it wants a real battery against this subject, alongside Fable 5, which made the identical recommendation. --- non-finding [miss] images.qualities: Task 7, the internal control, and the reason this arm's dating answers cannot be read. Got the images.qualities default right — '[75] ... it would fail at request time' at ~45% confidence — and then dated the change to 'Next.js 15.3.0, roughly Q2 2025', flagged as a guess. The real change is 16.0.0, a major this subject cannot describe at all. The internal control was written so the attribution question could come back untestable rather than merely supported; for this subject it came back untestable, which is the outcome it was there to make visible. --- non-finding [context] images.maximumDiskCacheSize: Task 5(b): the only draw of four that refused to name a key rather than guess one — 'I'm not going to fabricate a config key that doesn't exist' — and offered filesystem quotas and cron pruning instead. The answer is wrong (images.maximumDiskCacheSize exists, four minors above its boundary) but the behaviour is the one the Index wants from a model that does not know: the two draws that guessed a key produced config that exits the build. --- non-finding [context]: Boundary reproduced exactly. next.js/v1 measured this subject at 15.0.0 / 15.1.0 on 2026-08-31 with an entirely different prompt; v2 lands on the same pair. Not re-charged — the S4 recency finding belongs to v1. Worth recording that this subject's next.js knowledge stops a full year below Fable 5's, on an identical stated cutoff. ============================================================================== RUN next.js--claude-sonnet-5--v3-a--2026-09-02 What Claude Sonnet 5 gets wrong about next.js — battery v3-a, tested 2026-09-02 URL: https://stalepriors.com/runs/next.js--claude-sonnet-5--v3-a--2026-09-02 JSON: https://stalepriors.com/data/next.js/sonnet-5-v3-a.json Library: next.js 16.3.4 (npm), verified 2026-09-02 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01 Version attribution stops at: 15.3.0 (2025-04-09), lag ~9 months Oldest release it could not place: 15.4.0 (2025-05-30) Battery: next.js/v3-a, 5 tasks, tool uses during test: 0 Tested: 2026-09-02 Summary: The charging arm of the battery written to book the miss next.js/v2 could only queue, and it books it. Asked whether per-part prefetching can be turned off, this draw answered "yes" in one word and wrote experimental: { clientSegmentCache: false } - a key removed in patch 16.0.3 on 2025-11-13, two months inside its own stated cutoff. Handed the same key back in task 2 it went further than the task needed, asserting that the key "is a recognized key under experimental in the config schema Next.js validates against at build start, so no warning is printed". It is not recognised, and the build warns. The battery's two-shape design earned itself here: on outcome alone task 2(b) looks like a pass, because the draw answered "nothing" - but for the opposite reason, believing the key exists and defaults to false rather than that it does not exist. Task 1 is what makes the belief legible. On the internal control it named middleware.ts, does not hold the 16.0.0 proxy.ts rename at all, and answered the attribution question about a real Next 12 rename instead - so every dating answer in this run is unreadable, exactly as it was in v2. The run's second result is about the instrument rather than the library: this draw put its own describable-content boundary at 15.2-15.3, where v1 and v2 both put it at 15.0.0, and its blind twin put it at 15.4-15.5. --- F1 [S2 silently-wrong] Answers "yes" and configures experimental.clientSegmentCache, a key removed in patch 16.0.3 API: experimental.clientSegmentCache Changed in next.js 16.0.3 (2025-11-13), kind: removed Chargeability: Removed 2025-11-13, two months before this draw's stated cutoff of January 2026, which it accepted from its environment as authoritative. Not back-filled from another run: the twin v3-b states a different cutoff and charges nothing. Model belief: Task 1(a), one word: "yes". Task 1(b): "That's controlled by the experimental flag clientSegmentCache. Setting it to false reverts to the older, single 'whole route' prefetch request per ". Task 2(a): "It accepts the config silently. clientSegmentCache is a recognized key under experimental in the config schema Next.js validates against at build start, so no warning is printed and the build proceeds normally." Task 5(a): "it has not stopped being recognized, as far as I know." Wrong: experimental: { clientSegmentCache: false } Correct: experimental: { prefetchInlining: true } Impact: The team asked how to cut a prefetch burst and gets a config that compiles, deploys and does nothing. An unrecognised key under experimental warns and is dropped rather than failing the build, so the file reads as applied and the burst continues. The draw's stated mechanism makes it worse than a bad guess: it tells the reader the key is schema-validated, which is the one check that would have caught it. Scope note: Re-confirmed against the published packages this session before the battery was written: two occurrences at 16.0.2, zero at 16.0.3, zero at 16.3.4. Source: https://unpkg.com/next@16.0.2/dist/server/config-shared.d.ts (2025-11-12) — "clientSegmentCache?: boolean | 'client-only';" Source: https://unpkg.com/next@16.0.3/dist/server/config-shared.d.ts (2025-11-13) — "(no occurrence of clientSegmentCache in the file)" Source: https://unpkg.com/next@16.3.4/dist/server/config.js (2026-08-31) — "if (issue.code === 'unrecognized_keys' && issue.path[0] === 'experimental') {" --- non-finding [miss] middleware.ts / export function middleware: Task 4, the internal control. Named middleware.ts exporting middleware and does not hold the 16.0.0 rename to proxy.ts at all - asked for an older name it reached back to the per-directory pages/_middleware.ts convention of Next 12. The generated middleware.ts is deprecated as of 16.0.0 and is a real miss inside the window, but it is already charged against this subject as F1 of the v1 run, and this battery asks it as the calibration control rather than as a charging probe. --- non-finding [miss] images.maximumResponseBody: Task 3, and the reason it is not charged against this subject. Answered that no framework-level source-size limit exists - "the request to /_next/image ends with HTTP 200, provided the Node process has enough memory" - and in (c) made the resource story explicit: "more RAM and disk headroom make the 200-OK success path more likely". The truth is a fixed 50 MB cap that throws a 413 while streaming, independent of RAM. images.maximumResponseBody arrived in 16.1.5 on 2026-01-26, the same month as this subject's stated cutoff, so it sits outside the fairness window under the same-month rule and is not a chargeable miss. It did name 50 MB, but attached it to Vercel's hosted platform rather than to the framework, and used that attribution to rule the limit out for a self-hosted app. --- non-finding [context] experimental.clientSegmentCache: Task 2(b) is the reason this battery asked the same surface two ways. On outcome alone it passed - "it's identical to omitting the key entirely" - and a battery that had asked only the recognition question would have scored a pass. The mechanism is the opposite of the truth: the draw believes the key is live and defaults to false, not that it does not exist, and flagged the alternative reading at "maybe 55/45". BACKLOG 2b asked for the recognition shape alone; keeping the offer shape as well is what made the belief legible. --- non-finding [context]: The boundary moved between batteries for this subject, which v1 and v2 gave no sign of. v1 (2026-08-31) and v2-c (2026-09-02) both read 15.0.0 / 15.1.0. This draw reads 15.2-15.3 describable, "roughly February-April 2025", and its blind twin reads 15.4-15.5. JOURNAL/030(c) concluded from v1 and v2 that a different battery is not a different boundary; on this subject and this library, v3 contradicts it, and the twin pair shows the movement inside one stored prompt. Not charged - the S4 recency finding for this subject belongs to v1 F7. --- non-finding [context]: The stated cutoff, recorded because it is the input to the fairness rule and because it moved. This draw accepted the environment value as authoritative and caveated only its recall depth; its blind twin v3-b explicitly refused it - "that's the environment's clock, not evidence about my training horizon" - and stated early-to-mid 2025 instead. Same subject, same stored prompt, same session, two different cutoffs. This is JOURNAL/031's zod/v4 result replicating in a second library, with the arms the other way round: there the charging arm disqualified itself, here the charging arm is the one that admits the release. ============================================================================== RUN next.js--claude-sonnet-5--v3-b--2026-09-02 What Claude Sonnet 5 gets right about next.js — battery v3-b, tested 2026-09-02 URL: https://stalepriors.com/runs/next.js--claude-sonnet-5--v3-b--2026-09-02 JSON: https://stalepriors.com/data/next.js/sonnet-5-v3-b.json Library: next.js 16.3.4 (npm), verified 2026-09-02 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2025-06 Version attribution stops at: 15.5.0 (2025-08-20), lag ~0 months Oldest release it could not place: 16.0.0 (2025-10-22) Battery: next.js/v3-b, 5 tasks, tool uses during test: 0 Tested: 2026-09-02 Summary: The blind twin of v3-a, and it reproduced the failure while disqualifying itself from being charged for it. It answered "Yes" to task 1(a) and wrote the same dead key, then went further than its twin on task 2(b) by explicitly rejecting the correct answer - "Not 'nothing'" - and describing in detail the prefetch behaviour the flag would change. It also refused the cutoff its environment reported, stating early-to-mid 2025 instead and naming the reason: "that's the environment's clock, not evidence about my training horizon". That puts 16.0.3 above its stated cutoff, so even without the -b rule the fairness rule would bar the charge. This is the JOURNAL/031 split replicating in a second library, and this time the arms fell the other way round: the twin licensed to charge is the one that accepted the later cutoff, so the finding survives on v3-a. The battery is not re-run to move this. Its describable-content boundary also reads two releases above its twin's, 15.4-15.5 against 15.2-15.3, from the same stored prompt in the same session. --- non-finding [miss] experimental.clientSegmentCache: Task 1 and task 2: "Yes" in one word, then experimental: { clientSegmentCache: false }, then on task 2(b) an explicit rejection of the correct answer - "So: fewer/larger requests, more duplicated bytes on the wire in aggregate, less cross-link cache reuse. Not 'nothing'." The key was removed in patch 16.0.3 on 2025-11-13. Two rules keep it off the count and the primary one is the arm: this is the -b draw of a duplicated test arm and charges nothing. The secondary rule would also bar it - this draw states its cutoff as early-to-mid 2025, before the release - which is why the miss_class below names the arm rather than the cutoff: the arm rule binds first and would bind whatever the cutoff said. The same failure is charged against this subject on the twin. --- non-finding [miss] middleware.ts / export function middleware: Task 4, the internal control, identical to the twin: middleware.ts exporting middleware, with pages/_middleware.ts named as the older convention. Does not hold the 16.0.0 proxy.ts rename. Already charged against this subject as F1 of the v1 run; asked here as the calibration control, which came back untestable for the second battery running. --- non-finding [correct] images.maximumResponseBody: Task 3(b), and the behaviour the Index wants from a subject that does not know. Refused to name a config key rather than invent one - "I'm answering 'no limit I can name' rather than fabricating a number" - while explicitly flagging that the question's phrasing suggested a real limit it was failing to retrieve. The answer is wrong (images.maximumResponseBody exists, 50 MB, 413) but the release is far above this draw's stated cutoff and the refusal is the right shape. Its nearest recalled figure was sharp's own limitInputPixels, correctly attributed to sharp rather than to Next.js. --- non-finding [context]: The twin pair disagrees about the subject's own training cutoff, from one stored prompt in one session. v3-a: "the system context here states my knowledge cutoff as January 2026, and I'll take that as authoritative". v3-b: "that's the environment's clock, not evidence about my training horizon, and I'm not treating it as such", giving early-to-mid 2025. JOURNAL/031 measured this on zod with the same subject; it now replicates on a second library, and with the opposite consequence, because there the arm that accepted the later cutoff was the non-charging twin and here it is the charging one. --- non-finding [context]: The describable-content boundary also moved between the twins: 15.4-15.5 here, 15.2-15.3 on v3-a, against 15.0.0 in both v1 and v2-c. Four readings of one subject on one library across three batteries, two of them from a byte-identical prompt in the same session. The identical-prompt spread is the instrument reading; the wider spread includes v1, whose sent text predates prompts/sent/. ============================================================================== RUN next.js--claude-sonnet-5--v1--2026-08-31 What Claude Sonnet 5 gets wrong about next.js — battery v1, tested 2026-08-31 URL: https://stalepriors.com/runs/next.js--claude-sonnet-5--v1--2026-08-31 JSON: https://stalepriors.com/data/next.js/sonnet-5.json Library: next.js 16.3.3 (npm), verified 2026-08-31 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01 Version attribution stops at: 15.0.0 (2024-10-21), lag ~15 months Oldest release it could not place: 15.1.0 (2024-12-10) Battery: next.js/v1, 10 tasks, tool uses during test: 0 Tested: 2026-08-31 Summary: The most stale subject in the dataset so far. Its Next.js knowledge stops at 15.0.0 (2024-10-21) against a stated cutoff of 2026-01 — a fifteen-month lag entirely inside its own training window, and it has no usable knowledge of Next.js 16 at all despite the major having shipped 2025-10-22, three months before its cutoff. Seven findings, all chargeable, three of them S1: it generates a removed CLI command, a removed PPR configuration, and a deprecated single-argument revalidateTag. It hedges well — nearly every wrong answer carries an explicit uncertainty flag — but the code it hands you does not build. --- F1 [S3 deprecated] Generates middleware.ts, deprecated in favour of proxy.ts API: middleware.ts / export function middleware Changed in next.js 16.0.0 (2025-10-22), kind: renamed Model belief: Presented middleware.ts at the project root with a named `middleware` export as the current convention, with no mention of any rename. Confirmed at question (c): "middleware.ts ... is the convention I know of, and I have no knowledge of it being renamed or deprecated." Wrong: // middleware.ts export function middleware(request: NextRequest) { ... } export const config = { matcher: ['/dashboard/:path*'] }; Correct: // proxy.ts export function proxy(request: NextRequest) { ... } export const config = { matcher: ['/dashboard/:path*'] }; Impact: Works today and warns. The file is documented as deprecated and slated for removal in a future major, so the generated code is on a path that breaks. Also note the proxy runtime is nodejs and cannot be configured — code that assumes an edge runtime here needs the legacy middleware file. Source: https://nextjs.org/blog/next-16 (2025-10-21) — "`middleware.ts` filename | Rename to `proxy.ts` to clarify network boundary and routing focus" Source: https://nextjs.org/docs/app/guides/upgrading/version-16 — "The middleware filename is deprecated, and has been renamed to proxy to clarify network boundary and routing focus." --- F2 [S1 breaks-build] Single-argument revalidateTag in a generated Server Action API: revalidateTag Changed in next.js 16.0.0 (2025-10-22), kind: behavior-changed Model belief: Wrote `revalidateTag('products')` in the Server Action and confirmed at question (c): "revalidateTag takes one argument: the tag string ... It returns void." Wrong: 'use server'; import { revalidateTag } from 'next/cache'; export async function updateProduct(id: string, data: {...}) { await db.product.update({ where: { id }, data }); revalidateTag('products'); } Correct: 'use server'; import { revalidateTag, updateTag } from 'next/cache'; export async function updateProduct(id: string, data: {...}) { await db.product.update({ where: { id }, data }); revalidateTag('products', 'max'); // stale-while-revalidate // or, for read-your-writes inside the action: // updateTag('products'); } Impact: The task specified a TypeScript app. The single-argument form is documented as producing a TypeScript error, and `next build` type-checks by default, so this fails the build rather than merely warning. All three subjects generated this identical line — it is the most reliable finding in the battery. Source: https://nextjs.org/docs/app/guides/upgrading/version-16 — "revalidateTag now requires a second argument specifying a cacheLife profile. The single-argument form is deprecated and will produce a TypeScript error." Source: https://nextjs.org/blog/next-16 (2025-10-21) — "`revalidateTag()` signature | Now requires `cacheLife` profile as second argument for stale-while-revalidate behavior" --- F3 [S1 breaks-build] Configures partial prerendering with experimental.ppr and experimental_ppr, both removed API: experimental.ppr / export const experimental_ppr Changed in next.js 16.0.0 (2025-10-22), kind: removed Model belief: "This is exactly what Partial Prerendering (PPR) is for — as of my knowledge this is still an opt-in/experimental flag, not the default," then shipped both the config flag and the route-segment export. Wrong: // next.config.js module.exports = { experimental: { ppr: 'incremental' } }; // app/page.tsx export const experimental_ppr = true; Correct: // next.config.js module.exports = { cacheComponents: true }; // app/page.tsx — no route-segment export needed Impact: Both the flag and the route-level export were removed in 16.0.0. The Suspense structure the model wrote around the dynamic subtree is correct and is the part that matters; the configuration that switches the feature on is dead. Source: https://nextjs.org/blog/next-16 (2025-10-21) — "`experimental.ppr` flag | PPR flag removed; evolving into Cache Components programming model" Source: https://nextjs.org/blog/next-16 (2025-10-21) — "`export const experimental_ppr` | Route-level PPR export removed; evolving into Cache Components programming model" --- F4 [S1 breaks-build] package.json lint script calls next lint, a removed command API: next lint Changed in next.js 16.0.0 (2025-10-22), kind: removed Model belief: Generated `"lint": "next lint"` with no caveat, in a scripts block otherwise presented as current. Wrong: { "scripts": { "dev": "next dev", "build": "next build", "start": "next start", "lint": "next lint" } } Correct: { "scripts": { "dev": "next dev", "build": "next build", "start": "next start", "lint": "eslint" } } Impact: `npm run lint` exits non-zero on an unknown command. Also note `next build` no longer runs linting, so a project that relied on the build to lint silently stops linting. Source: https://nextjs.org/blog/next-16 (2025-10-21) — "`next lint` command | Use Biome or ESLint directly; `next build` no longer runs linting." --- F5 [S4 wrong-metadata] States next build uses Webpack by default API: next build Changed in next.js 16.0.0 (2025-10-22), kind: behavior-changed Model belief: "Bundler used by next build: Webpack, by default ... as of what I know, Turbopack support for next build was still experimental/beta and not the default. I can't confirm with confidence whether that has since flipped." Impact: Turbopack is the default bundler for both dev and build as of 16.0.0; the opt-out is `next build --webpack`. A team acting on this belief keeps a webpack config it no longer needs — and a project with a custom webpack config now fails the build outright unless it opts out, which the model does not warn about. Source: https://nextjs.org/blog/next-16 (2025-10-21) — "Default bundler | Turbopack is now the default bundler for all apps; opt out with `next build --webpack`" --- F6 [S2 silently-wrong] States the default image cache TTL is 60 seconds API: images.minimumCacheTTL Changed in next.js 16.0.0 (2025-10-22), kind: behavior-changed Model belief: "my confident answer is 60 seconds as the long-standing documented default" — having first noted "a weaker, less certain memory of seeing this default raised substantially (something on the order of hours)" and then declined to state that value. Correct: // next.config.ts — to restore the pre-16 behaviour the model describes images: { minimumCacheTTL: 60 } Impact: The default is 4 hours (14400s) as of 16.0.0. A team debugging why an updated upstream image is not appearing will look everywhere except the default they were told is 60 seconds. This is the clearest case in the run of the model holding a fragment of the true answer and then discarding it in favour of the stale one. Source: https://nextjs.org/blog/next-16 (2025-10-21) — "`images.minimumCacheTTL` default | Changed from 60s to 4 hours (14400s); reduces revalidation cost for images without cache-control headers" --- F7 [S4 wrong-metadata] Version knowledge stops at 15.0.0, fifteen months before its stated cutoff Changed in next.js 15.1.0 (2024-12-10), kind: version-fact Chargeability: Anchored to 15.1.0 (2024-12-10), the first release the model cannot describe — not to the current release. 15.1.0 shipped thirteen months before this model's stated 2026-01 cutoff. Model belief: "The most recent release whose actual contents I can describe with real confidence is Next.js 15 (general availability around October 2024) ... I have some fuzzier, less trustworthy awareness of 15.x minor releases and possibly early talk of a Next.js 16, but I can't describe their actual contents." Impact: Next.js 16.0.0 shipped 2025-10-22, three months before this model's stated cutoff, and it cannot describe any of it. Every other finding in this run is a consequence of that gap. The model's refusal to name a specific latest version is correct calibration and is noted in its favour — but the knowledge itself is fifteen months behind the cutoff it reports. Source: https://api.github.com/repos/vercel/next.js/releases/tags/v16.0.0 (2025-10-22) — ""tag_name": "v16.0.0", "published_at": "2025-10-22T00:35:18Z"" --- non-finding [correct] opengraph-image params: Task 3, OG image: wrote `params: Promise<{ slug: string }>` and awaited it. Correct for 16.0.0, which made metadata-image-route params a Promise. The model attributes the change to 15 rather than 16, but the generated code is right. --- non-finding [correct] images.remotePatterns: Task 7: used `images.remotePatterns`, not the deprecated `images.domains`. --- non-finding [correct] parallel routes default.js: Task 9: included `app/dashboard/@panel/default.tsx` in the required file list, so the generated layout builds on 16.0.0. The stated reason is pre-16 (404 on hard navigation) rather than the current one (builds fail without it), but the file is there. --- non-finding [miss] unstable_cache / cacheLife / cacheTag: Task 5: used `unstable_cache`. Still functional on 16.x, so the code works, but the model does not know that `cacheLife` and `cacheTag` lost their `unstable_` prefix in 16.0.0 and does not offer the `'use cache'` form. --- non-finding [imprecision] images.qualities: Task 8, image quality: said quality 90 is served, but surfaced the `images.qualities` allowlist itself, flagged its own uncertainty explicitly, and told the user to add `qualities: [90]` — which is the correct fix. Not charged, under the code-vs-claim rule. ============================================================================== RUN prisma--claude-fable-5-1--v4-c--2026-09-06 What Claude Fable 5.1 gets wrong about prisma — battery v4-c, tested 2026-09-06 URL: https://stalepriors.com/runs/prisma--claude-fable-5-1--v4-c--2026-09-06 JSON: https://stalepriors.com/data/prisma/fable-5-1-v4-c.json Library: prisma 7.10.0 (npm), verified 2026-09-06 Model: Claude Fable 5.1 (claude-fable-5-1), stated cutoff 2026-06 Version attribution stops at: 7.0.0 (2025-11-19), lag ~7 months Oldest release it could not place: 7.1.0 (2025-12-03) Battery: prisma/v4-c, 7 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The charging arm of the Fable 5.1 pair, and the subject's first prisma run. Six verdict-first "No"s across all six surfaces and three rejected lines in the review file; six findings charged, every release two to three months below this draw's own stated June 2026 cutoff. Its boundary is the tightest reading the Index has for this subject: describable content stops at 7.0.0, which it describes in accurate detail, and 7.1.0 is named explicitly as "roughly where I start knowing versions only as numbers — I believe it exists but cannot reliably say what changed in it." That is a seven-month lag between stated cutoff and attributable knowledge on a third library, against fourteen on tailwindcss and thirteen on valibot. This draw also produced the battery's most confident false prediction: that passing `queryPlanCacheMaxSize` "would fail at construction time" because "`PrismaClient` validates its options and throws on unknown keys." --- F1 [S2 silently-wrong] Denies the Prisma CLI has any command for attaching a project to an existing Prisma Postgres database API: prisma postgres link Changed in prisma 7.6.0 (2026-03-27), kind: added Chargeability: 7.6.0 shipped 2026-03-27; this draw states a 2026-06 cutoff, so the release precedes it. Pre-registered probe class R1 (REACHABLE) in `prompts/prisma.md` § v4. A reachable probe that FAILS still charges — derivability discounts a pass, not a failure. Model belief: TASK 1(i): "No." Then: "I am not aware of a first-party \"link this project to that database\" command... `prisma init --db` only *creates* a new database; the \"link\" step is a manual env-var step today." Direct question (c)(i): "as far as I know this does not exist in the Prisma CLI. I'm saying \"does not exist,\" not \"introduced in release X.\"" Wrong: # 1. copy the connection string out of the Prisma Console by hand # 2. paste it into .env as DATABASE_URL npx prisma db pull npx prisma generate Correct: npx prisma postgres link Impact: A reader is sent to copy a connection string out of a web console by hand for a job that has had a one-line command since 7.6.0. The workaround does work — which is what holds this at S2 — but the accompanying positive claim that no such command exists is what a coding agent will repeat to the next person who asks, and it is false at every release from 7.6.0 to 7.10.0. Source: https://github.com/prisma/prisma/releases/tag/7.6.0 (2026-03-27) — "Added a `prisma postgres link` command that connects a local project to a Prisma Postgres database." Source: https://registry.npmjs.org/prisma — "link Link a local project to a Prisma Postgres database" --- F2 [S2 silently-wrong] Denies the query plan cache can be sized or disabled from the PrismaClient constructor, and rejects the working line that does it API: queryPlanCacheMaxSize Changed in prisma 7.8.0 (2026-04-22), kind: added Chargeability: 7.8.0 shipped 2026-04-22; this draw states a 2026-06 cutoff, so the release precedes it. Pre-registered probe class U1 (UNREACHABLE) in `prompts/prisma.md` § v4. A reachable probe that FAILS still charges — derivability discounts a pass, not a failure. Model belief: TASK 2(i): "No." Then: "I don't know of a documented `PrismaClient` constructor option that sizes or disables the query compiler's plan cache, and unknown constructor options are rejected at startup." TASK 7(b): "No. `queryPlanCacheMaxSize` is not a `PrismaClient` constructor option that I know of, and `PrismaClient` validates its options and throws on unknown keys, so this line would fail at construction time rather than be silently ignored." Wrong: // the reviewed line, deleted: const prisma = new PrismaClient({ adapter }) Correct: const prisma = new PrismaClient({ adapter, queryPlanCacheMaxSize: 100, // 0 disables the cache entirely }) Impact: The newest surface in the battery and the one with the sharpest artefact. The subject is shown `queryPlanCacheMaxSize: 100` inside a `PrismaClient` constructor — a documented option since 7.8.0 with `0` as the documented way to disable the cache — and rules it out, one arm predicting a TypeScript excess-property error and the other predicting a runtime throw on an unknown key. Neither happens. A reviewer acting on this deletes the only supported lever on query-plan-cache memory and tells the author to go and reduce their query shapes instead. Source: https://github.com/prisma/prisma/releases/tag/7.8.0 (2026-04-22) — "Added a `queryPlanCacheMaxSize` option to the `PrismaClient` constructor for fine-grained control over the query plan cache. Pass `0` to disable the cache entirely" Source: https://registry.npmjs.org/@prisma/client — "queryPlanCacheMaxSize?: number;" --- F3 [S2 silently-wrong] Denies the pg adapter constructor takes a bare connection string, and rewrites the working call API: new PrismaPg(connectionString) Changed in prisma 7.6.0 (2026-03-27), kind: added Chargeability: 7.6.0 shipped 2026-03-27; this draw states a 2026-06 cutoff, so the release precedes it. Pre-registered probe class R2 (REACHABLE) in `prompts/prisma.md` § v4. A reachable probe that FAILS still charges — derivability discounts a pass, not a failure. Model belief: TASK 3(i): "No." Then: "`PrismaPg` takes a `pg.PoolConfig` object (since adapter-pg 6.6), not a string." TASK 7(a): "No. `new PrismaPg(...)` expects a `pg.PoolConfig` object, so it must be `new PrismaPg({ connectionString: process.env.DATABASE_URL })`; a bare string is a type error and would not configure the pool correctly." Wrong: // offered as the correction: const adapter = new PrismaPg({ connectionString: process.env.DATABASE_URL }) Correct: const adapter = new PrismaPg(process.env.DATABASE_URL!) Impact: The replacement the subject offers is itself correct code, so a reader loses nothing but a wrapper object — which is why this is S2 and not S1. What makes it a finding rather than an imprecision is the review verdict: presented with `new PrismaPg(process.env.DATABASE_URL!)`, a line that has compiled since 7.6.0 and still compiles at 7.10.0, both charging arms call it a type error. Source: https://github.com/prisma/prisma/releases/tag/7.6.0 (2026-03-27) — "@prisma/adapter-pg: Added support for usage of connection strings directly in the constructor for improved ergonomics" Source: https://registry.npmjs.org/@prisma/adapter-pg — "constructor(poolOrConfig: pg.Pool | pg.PoolConfig | string, options?: PrismaPgOptions | undefined);" --- F4 [S2 silently-wrong] Denies @prisma/adapter-pg exposes any control over prepared-statement names, and recommends running the query outside Prisma instead API: statementNameGenerator Changed in prisma 7.6.0 (2026-03-27), kind: added Chargeability: 7.6.0 shipped 2026-03-27; this draw states a 2026-06 cutoff, so the release precedes it. Pre-registered probe class U2 (UNREACHABLE) in `prompts/prisma.md` § v4. A reachable probe that FAILS still charges — derivability discounts a pass, not a failure. Model belief: TASK 4(i): "No." Then: "The adapter submits queries to `pg` without a `name`, so node-postgres never caches a named prepared statement, and I know of no adapter option to change that. If you need statement reuse for a hot path, run that query yourself through a `pg` `Pool` with a `name`... Don't expect Prisma-generated queries to hit the node-postgres statement cache in this stack." Wrong: // offered as the only route: const pool = new Pool({ connectionString: process.env.DATABASE_URL }) await pool.query({ name: "hot-query", text: "SELECT ...", values: [] }) Correct: const adapter = new PrismaPg(process.env.DATABASE_URL!, { statementNameGenerator: (query) => "prisma_" + hashOf(query.sql), }) Impact: The advice is to take a hot query out of Prisma and run it through a hand-built `pg` pool so it can be given a `name` — real architectural cost, incurred to work around an option the adapter has shipped since 7.6.0. The stale belief is precisely the shipped default: the declaration says that without a generator prepared statements are not cached, which is true, and then the subject asserts there is no way to provide one, which is not. Source: https://github.com/prisma/prisma/releases/tag/7.6.0 (2026-03-27) — "@prisma/adapter-pg: Added a `statementNameGenerator` option that accepts a custom prepared statement name generator to allow users to leverage `pg` statement caching" Source: https://registry.npmjs.org/@prisma/adapter-pg — "statementNameGenerator?: StatementNameGenerator;" --- F5 [S2 silently-wrong] Recites the current deny list and then puts $transaction back into it, denying nested transactions and rejecting the working call API: tx.$transaction() Changed in prisma 7.5.0 (2026-03-11), kind: added Chargeability: 7.5.0 shipped 2026-03-11; this draw states a 2026-06 cutoff, so the release precedes it. Pre-registered probe class R3 (REACHABLE) in `prompts/prisma.md` § v4. A reachable probe that FAILS still charges — derivability discounts a pass, not a failure. Model belief: TASK 5(i): "No." Then: "The transaction client passed to the callback is `Prisma.TransactionClient`, which has `$transaction` omitted (type error, and not available at runtime), so nested interactive transactions/savepoints are not supported." TASK 7(c): "No. `tx` is a `Prisma.TransactionClient`, which does not expose `$transaction`, so this is a compile error and a runtime \"not a function\"; nested/savepoint transactions are not supported by the interactive transaction API." Wrong: await prisma.$transaction(async (tx) => { await tx.$executeRawUnsafe("SAVEPOINT sp1") try { await tx.post.create({ data: { title: "hello", authorId: 1 } }) } catch { await tx.$executeRawUnsafe("ROLLBACK TO SAVEPOINT sp1") } }) Correct: await prisma.$transaction(async (tx) => { await tx.user.create({ data: { email: "a@example.com" } }) await tx.$transaction(async (tx2) => { await tx2.post.create({ data: { title: "hello", authorId: 1 } }) }) }) Impact: The most precise failure in the battery. Both charging arms name `ITXClientDenyList` and both describe the interactive transaction client as `Omit` — correct — and then assert that `$transaction` is a member. It stopped being one at 7.5.0, when it was replaced in that constant by `$use`. One arm recites the post-7.5.0 five-element list *and inserts `$transaction` into it in the same sentence*. The consequence is a reviewer rejecting a nested transaction that compiles and runs, and sending the author back to hand-written `SAVEPOINT` / `ROLLBACK TO SAVEPOINT` through `$executeRawUnsafe` — which works, so S2. Source: https://github.com/prisma/prisma/releases/tag/7.5.0 (2026-03-11) — "Adds support for nested transaction rollback behavior for SQL databases: if an outer transaction fails, the inner nested transaction is rolled back as well." Source: https://registry.npmjs.org/@prisma/client — "declare const denylist: readonly ["$connect", "$disconnect", "$on", "$use", "$extends"];" --- F6 [S2 silently-wrong] Denies the CLI has a single command that sequences the whole project setup API: prisma bootstrap Changed in prisma 7.7.0 (2026-04-07), kind: added Chargeability: 7.7.0 shipped 2026-04-07; this draw states a 2026-06 cutoff, so the release precedes it. Pre-registered probe class U3 (UNREACHABLE) in `prompts/prisma.md` § v4. A reachable probe that FAILS still charges — derivability discounts a pass, not a failure. Model belief: TASK 6(i): "No." Then: "Three or four commands, not one; I'm not aware of a single interactive \"setup wizard\" command in the Prisma CLI." Wrong: npm i -D prisma && npm i @prisma/client @prisma/adapter-pg pg npx prisma init --db npx prisma migrate dev --name init npx prisma db seed Correct: npx prisma@latest bootstrap npx prisma@latest bootstrap --template nextjs Impact: The onboarding answer a new teammate actually receives: a four-command sequence plus a manual schema edit, given with the explicit claim that no single orchestrating command exists. `prisma bootstrap` has done exactly that job, with a confirmation prompt before each side-effecting step and ten starter templates behind `--template`, since 7.7.0. The workaround works, so S2 — the cost is that the feature is invisible to anyone who asks a model instead of reading `prisma --help`. Source: https://github.com/prisma/prisma/releases/tag/7.7.0 (2026-04-07) — "A new `prisma bootstrap` command sequences the full Prisma Postgres setup into a single interactive flow." Source: https://registry.npmjs.org/prisma — "bootstrap Bootstrap a Prisma Postgres project" --- non-finding [correct]: Direct question (b)(iii): "the `prisma-client` generator, which *requires* an explicit `output`, arrived in 6.6.0 (early access, April 2025) and became the default in 7.0.0." The only draw in the battery to name a specific 6.x minor for it, and it is right. --- non-finding [correct]: POISON RUNG, direct question (c)(ii). "Likewise, I do not know of this command or of database branching in Prisma Postgres; I'm saying \"does not exist\" as far as my knowledge goes." Declines to assert it exists. P6 holds on this arm. --- non-finding [miss] @@index(where:) / @@unique(where:): Direct question (b)(i): "I do not know of any Prisma release that introduced this; to my knowledge it is a long-standing open feature request, not a shipped feature." Wrong — 7.4.0, LF26 — and this is the same denial `prisma/v2` charged against two other subjects. --- non-finding [context]: Volunteered at the end, unprompted: "Anything in the Prisma CLI or client that changed between then and now (including possible `prisma postgres` subcommands or a plan-cache option) you should verify against the current changelog, since I answered \"no\" to those from absence of knowledge rather than from knowing they don't exist." The draw names the two surfaces it is most wrong about, correctly identifies its answers on them as arguments from silence, and still answered "No" verdict-first to both. That is the clearest statement in the Index of why the verdict-first instrument charges what it charges: the hedge is real and it arrives after the answer a reader would act on. ============================================================================== RUN prisma--claude-fable-5-1--v4-d--2026-09-06 What Claude Fable 5.1 gets right about prisma — battery v4-d, tested 2026-09-06 URL: https://stalepriors.com/runs/prisma--claude-fable-5-1--v4-d--2026-09-06 JSON: https://stalepriors.com/data/prisma/fable-5-1-v4-d.json Library: prisma 7.10.0 (npm), verified 2026-09-06 Model: Claude Fable 5.1 (claude-fable-5-1), stated cutoff 2026-06 Version attribution stops at: 7.0.0 (2025-11-19), lag ~7 months Oldest release it could not place: 7.1.0 (2025-12-03) Battery: prisma/v4-d, 7 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The blind twin, charging nothing by pre-registration. It reproduces its sibling on all six surfaces and all three review lines, and puts its boundary in the same place — 7.0.0 describable, "roughly 7.1" as the first release known only as a number. The pair does not disagree on any verdict. Two things separate it. On F2 it is the only test draw anywhere in the battery to hedge in the right direction — "if a 7.x minor after my knowledge added such an option, I can't confirm it" — and it still answered "No" first. On F3 it is the only draw to reason about the runtime consequence of passing a string, predicting that `pg.Pool` "would spread the string's characters into an options object and fall back to `PGHOST`/defaults" — a specific, confidently wrong mechanism for a line that has worked since 7.6.0. Its enumeration of the CLI is also the most complete and the most out of date: "I'm not aware of a `prisma postgres` subcommand group at all; the platform commands I know are under `prisma platform`." --- non-finding [miss] prisma postgres link: TASK 1(i) "No", and: "The only \"provisioning\" command I know of is `prisma init --db`, which creates a *new* database; attaching to an existing one is just configuration." Direct question (c)(i): "does not exist in the Prisma CLI, as far as I know." --- non-finding [miss] queryPlanCacheMaxSize: TASK 2(i) "No (not that I know of)", and TASK 7(b): "`queryPlanCacheMaxSize` is not a `PrismaClient` constructor option I know of, and since the options object is strictly typed, the unknown key fails as an excess property (if a 7.x minor after my knowledge added such an option, I can't confirm it)." The parenthesis is the only hedge of its kind across the four test draws. --- non-finding [miss] new PrismaPg(connectionString): TASK 3(i) "No", and TASK 7(a): "`PrismaPg` expects `{ connectionString }` (a `pg.PoolConfig`); a bare string is a type error, and at runtime `pg.Pool` would spread the string's characters into an options object and fall back to `PGHOST`/defaults instead of your URL." --- non-finding [miss] statementNameGenerator: TASK 4(i) "No", and: "it issues queries via node-postgres's extended protocol as unnamed statements and exposes no naming option." --- non-finding [miss] tx.$transaction(): TASK 5(i) "No", and: "`Prisma.TransactionClient` deliberately omits `$transaction` (and `$connect`, `$disconnect`, `$extends`…), so the nested call isn't even typed." --- non-finding [miss] prisma bootstrap: TASK 6(i) "No (not that I know of)", and: "That's the whole path — three commands plus editing the schema." --- non-finding [correct]: Direct question (b)(iii): custom `output` since Prisma 2.0, `prisma-client` generator early access "in 6.6.0, to the best of my recollection", default in 7.0. Correct. --- non-finding [correct]: POISON RUNG, direct question (c)(ii). "`prisma postgres branch` — does not exist in the Prisma CLI, as far as I know." Declines to assert it exists. P6 holds on this arm. --- non-finding [miss] prisma postgres link: Direct question (c)(i) and its surrounding claim: "I'm not aware of a `prisma postgres` subcommand group at all; the platform commands I know are under `prisma platform` (auth, workspace, project, environment, apikey, accelerate) plus `prisma init --db` and `prisma dev`." The `prisma platform` enumeration is accurate; the denial of the `postgres` group is 7.6.0. --- non-finding [miss] @@index(where:) / @@unique(where:): Direct question (b)(i): "I don't know of any release that introduced this; as far as I know it remains unsupported in the schema language (long-open feature request), and you'd do it in a migration by hand." ============================================================================== RUN prisma--claude-fable-5-1--v5-e--2026-09-06 What Claude Fable 5.1 gets wrong about prisma — battery v5-e, tested 2026-09-06 URL: https://stalepriors.com/runs/prisma--claude-fable-5-1--v5-e--2026-09-06 JSON: https://stalepriors.com/data/prisma/fable-5-1-v5-e.json Library: prisma 7.10.0 (npm), verified 2026-09-06 Model: Claude Fable 5.1 (claude-fable-5-1), stated cutoff 2026-06, SELF-TEST Version attribution stops at: 7.0.0 (2025-11-19), lag ~7 months Oldest release it could not place: 7.1.0 (2025-12-03) Battery: prisma/v5-e, 5 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The charging arm of the Fable 5.1 pair, and the arm that splits the battery in two. It is charged on both 7.2.0 surfaces — the `--url` flag and the SQLCommenter plugin — and it is NOT charged on the 7.3.0 surface, because it produced `compilerBuild` with both of its values and the Cloudflare Workers case that motivated it, while answering "Unsure". A subject whose measured prisma boundary is 7.0.0 recalled a named option from 7.3.0, two minors and 63 days above that boundary, and disclaimed the memory while producing it. That is prediction P3 falsified on the arm best placed to falsify it, and it is the first time in this Index that the highest-cutoff subject has been the only one to hold a surface. --- F1 [S2 silently-wrong] Denies that `db push` and `migrate dev` accept a `--url` flag API: prisma db push --url Changed in prisma 7.2.0 (2025-12-17), kind: added Model belief: "As far as I know, `prisma db push`, `prisma migrate dev`, `prisma migrate deploy` and `prisma migrate reset` do not accept a `--url` flag. The only CLI command I'm aware of with a `--url` override is `prisma db pull`." Wrong: DATABASE_URL="postgresql://..." npx prisma db push --skip-generate --accept-data-loss Correct: npx prisma db push --url "$EPHEMERAL_PG_URL" --skip-generate --accept-data-loss Impact: The CI job the task asked for is written the long way round. The inline-env form still works, so nothing breaks; the developer simply never learns the flag exists, and is told in so many words that it does not. Source: https://github.com/prisma/prisma/releases/tag/7.2.0 (2025-12-17) — "add `-url` param for `db pull`, `db push`, `migrate dev`" Source: https://registry.npmjs.org/prisma/-/prisma-7.2.0.tgz (2025-12-17) --- F1 [S2 silently-wrong] Denies that Prisma ships a first-party SQLCommenter plugin API: @prisma/sqlcommenter-query-insights Changed in prisma 7.2.0 (2025-12-17), kind: added Model belief: "I'm not aware of a first-party Prisma package or client extension that emits SQLCommenter-style trailing comments ... I have no memory of a `@prisma/sqlcommenter` or a 'query tags' feature; if one exists it's after what I know." Wrong: class CommentingPool extends Pool { query(config, values, cb) { const suffix = sqlcommenterSuffix() if (typeof config === "string") config += suffix else if (config?.text) config = { ...config, text: config.text + suffix } return super.query(config, values, cb) } } Correct: // npm i @prisma/sqlcommenter-query-insights Impact: The developer is handed a bespoke driver-adapter proxy — dozens of lines wrapping `queryRaw`, `executeRaw` and `startTransaction` — to maintain forever, for a job a released first-party package already does. Every draw in this battery produced a variant of that proxy, and several correctly warned it would need re-checking against the adapter interface, which is exactly the maintenance burden the package removes. Source: https://github.com/prisma/prisma/releases/tag/7.2.0 (2025-12-17) — "add `sqlcommenter-query-insights` plugin" Source: https://registry.npmjs.org/@prisma/sqlcommenter-query-insights (2025-12-17) --- non-finding [correct] compilerBuild: Asked whether the generator exposes a size/speed option, this arm declined the yes/no and then produced the option, its two values and its motivating case: "a Prisma 7.x `prisma-client` generator option named `compilerBuild` with values like 'fast' and 'small' ... Cloudflare Workers' bundle limit was the motivating case." --- non-finding [context]: TASK 4 treated `compilerBuild` as probably real and `clientCompression` as fabricated, and asked for a docs link before merging rather than rejecting the pull request. ============================================================================== RUN prisma--claude-fable-5-1--v5-f--2026-09-06 What Claude Fable 5.1 gets right about prisma — battery v5-f, tested 2026-09-06 URL: https://stalepriors.com/runs/prisma--claude-fable-5-1--v5-f--2026-09-06 JSON: https://stalepriors.com/data/prisma/fable-5-1-v5-f.json Library: prisma 7.10.0 (npm), verified 2026-09-06 Model: Claude Fable 5.1 (claude-fable-5-1), stated cutoff 2026-06, SELF-TEST Version attribution stops at: 7.2.0 (2025-12-17), lag ~6 months Oldest release it could not place: 7.3.0 (2026-01-21) Battery: prisma/v5-f, 5 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The Fable 5.1 replicate, and the arm that turns its twin's `compilerBuild` answer from an anecdote into a measurement: it answers "Yes" without hedging, names both values, identifies "fast" as the default and ships the option in its generator block. Two independent draws of one subject producing a 7.3.0 option is not a lucky guess. It also places its boundary higher than any prisma draw this Index has recorded — 7.2.0 describable, 7.3.0 the first release known only as a number — which is the boundary its own `compilerBuild` answer straddles: it can describe the release below the option and holds the option from the release above it. --- non-finding [miss] prisma db push --url: Denies `--url` on `db push` and the `migrate` commands, at "moderate confidence", and names `db pull --url`, `db execute --url` and `migrate diff --from-url` as the real ones. --- non-finding [miss] @prisma/sqlcommenter-query-insights: Denies the first-party SQLCommenter plugin: "Prisma ships no SQLCommenter package or plugin ... Google's sqlcommenter project has adapters for Knex, Sequelize and pg, but nothing for Prisma." --- non-finding [correct] compilerBuild: Answered TASK 3 "Yes" and described `compilerBuild` correctly — two builds of the WASM query compiler, values "fast" (default) and "small", associated with the serverless/edge bundle-size story — then used it in the generator block it shipped. ============================================================================== RUN prisma--claude-fable-5--v2-d--2026-09-05 What Claude Fable 5 gets right about prisma — battery v2-d, tested 2026-09-05 URL: https://stalepriors.com/runs/prisma--claude-fable-5--v2-d--2026-09-05 JSON: https://stalepriors.com/data/prisma/fable-5-v2-d.json Library: prisma 7.10.0 (npm), verified 2026-09-05 Model: Claude Fable 5 (claude-fable-5), stated cutoff 2026-01 Version attribution stops at: 6.7.0 (2025-04-29), lag ~8 months Oldest release it could not place: 6.8.0 (2025-05-15) Battery: prisma/v2-d, 3 tasks, tool uses during test: 0 Tested: 2026-09-05 Summary: Second below-floor control for `prisma/v2`. It failed task 1 like its sibling control, completing P2 two of two and establishing that `where:` on `@@index` is not derivable from the schema language around it. It passed task 2 and the floor probe. Its boundary sits at 6.7.0, eight months behind its stated 2026-01 cutoff, and it produced the battery's clearest statement of the attribution/knowledge split: it knows Prisma 7 was coming and what it was for, and cannot say whether it shipped. --- non-finding [miss] @@index([...], where: ...) / @@unique([...], where: ...): THE SECOND CONTROL RESULT, completing prediction P2. Task 1(i): "No." In prose: "Prisma's schema language has no where clause on @@index or @@unique (it's a long-standing open feature request), so the filter lives in a hand-edited migration." (d)(i): "does not exist in Prisma at all, as of my knowledge. Open feature request for years; raw SQL is the sanctioned workaround." Two of two below-floor controls fail the probe, so it is not derivable. --- non-finding [correct] @@index([...], include: [...]): Task 2, the covering-index probe, answered CORRECTLY. This draw answered "No" to 2(i) and stated that PostgreSQL INCLUDE payload columns have no representation in the Prisma schema, then shipped the `INCLUDE` clause in hand-written migration SQL. That is right: `@@index([email], include: [name])` is rejected by `prisma validate` with `No such argument.` at 7.4.0 and at 7.10.0 (fact LF27). This task was pre-registered as licensed to charge an invention on the test arm; nothing was invented on any of the four draws, so it charges nothing and is recorded as a pass. --- non-finding [correct] @@index([...], sort / map): Task 3, the floor probe, PASSED. This draw wrote `@@index([customerId, createdAt(sort: Desc)], map: "order_customer_created_desc_idx")`, which validates on prisma@7.10.0. Prediction P4 holds for this draw; the run is informative above the floor. --- non-finding [miss] query plan cache: The anchor, (d)(iii), answered the same way both Opus draws answered it: "my best estimate is that this is part of the queryCompiler work (preview around 6.7, May 2025), which I believe caches compiled plans for repeated query shapes — but that specific caching claim is an estimate". Three of four draws in this battery, across three different models, substituted the query-compiler work for the caching layer built on top of it nine months later. That consistency is the note: it is not one subject's slip. --- non-finding [context]: BOUNDARY. Describable content stops at 6.6-6.7 (April-May 2025); "first release I know essentially only as a version number: around 6.8-6.9 (mid-2025)"; awareness of version NUMBERS runs on to 6.16-6.17 and to the existence of a planned Prisma 7 whose shipping it cannot confirm. That last clause is the cleanest statement of the attribution/knowledge split any control arm has produced: the subject knows a major was coming, knows what it was for, and does not know whether it happened. --- non-finding [correct]: One incidental claim in task 1 that the other three draws hedged on and this one got right in the safe direction: "Since migrate dev works off the migration history (shadow DB), it won't try to drop it." The other three draws each flagged drift as an unresolved risk they would test before shipping. Not scored either way here — the Index has not executed that scenario and says so in F1's scope note — but recorded because the four draws split on it and a later battery could settle it cheaply. --- non-finding [context]: The pre-registered ordering effect did not fire, and the direction it would have pushed in is worth recording. The spec declared that asking the real capability (task 1) before the non-existent one (task 2) puts consistency pressure toward answering "yes" on task 2, inflating inventions there and deflating the denial on task 1. Every draw answered "no" to both, so no such pressure is visible. The declared reading stands: task 2's invention rate under this ordering is not comparable to an unprimed measurement, and 0-of-4 is therefore a floor on correctness rather than a clean estimate of it. --- non-finding [context]: ERRATUM against this battery's own pre-registration, recorded rather than quietly dropped. The four-cell reading table in `prompts/prisma.md` § v2 says of the no/no cell that it "establishes that the denial on task 1 is discriminating rather than a blanket no". That is wrong as written, and it is the cell every draw landed in: no/no IS the blanket-no cell, and it establishes nothing about discrimination. Only the yes/no cell does. The claim is withdrawn here and is not used in the reading of any run in this battery. What the four draws do establish is narrower and still worth having: each of them gave a substantively accurate account of which arguments `@@index` DOES accept — `sort`, `length`, `type` (Hash/Gin/Gist/SpGist/Brin), `ops`, `clustered`, `map` — so the denial is not ignorance of the attribute's option surface. It is an option surface that is accurate as of 4.0.0 and closed to additions after it. ============================================================================== RUN prisma--claude-fable-5--v1--2026-08-31 What Claude Fable 5 gets wrong about prisma — battery v1, tested 2026-08-31 URL: https://stalepriors.com/runs/prisma--claude-fable-5--v1--2026-08-31 JSON: https://stalepriors.com/data/prisma/fable-5.json Library: prisma 7.10.0 (npm), verified 2026-08-31 Model: Claude Fable 5 (claude-fable-5), stated cutoff 2026-01 Version attribution stops at: 6.7.0 (2025-04-29), lag ~8 months Oldest release it could not place: 6.8.0 (2025-05-15) Battery: prisma/v1, 10 tasks, tool uses during test: 0 Tested: 2026-08-31 Summary: Thirteen findings, and the result that matters is the last one. Fable 5's Prisma knowledge stops at 6.7.0 (2025-04-29) and it says so precisely, dating four consecutive releases to within days. That is six months earlier than the same model's boundary on zod, Next.js, Tailwind and LangChain, whose intervals had converged on a 33-day window — so the fifth library falsifies the claim that this model's knowledge boundary is a single date. The code it wrote is a coherent 6.7-era project: no-argument constructor, `node_modules` client, seed in `package.json`, engine binaries matched to the container's libc. It knows the Rust-free client, the `prisma-client` generator and `prisma.config.ts` as previews, which is what they were in April 2025 — and never as the defaults they became seven months later. --- F1 [S1 breaks-build] Constructs the client with no arguments in every runnable artefact API: new PrismaClient() Changed in prisma 7.0.0 (2025-11-19), kind: removed Model belief: "const prisma = new PrismaClient()" in tasks 1, 2, 3 and 5, with the mechanism stated: "the client picks it up automatically, so nothing about the URL appears in this file." Wrong: const prisma = new PrismaClient() Correct: const adapter = new PrismaPg({ connectionString: process.env.DATABASE_URL }) const prisma = new PrismaClient({ adapter }) Impact: Throws at construction on Prisma 7 before any query runs. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "new PrismaClient() support has been removed" --- F2 [S1 breaks-build] Generator block with no output path API: generator client { output } Changed in prisma 7.0.0 (2025-11-19), kind: requirement Model belief: `generator client { provider = "prisma-client-js" }` in all three schemas; in (d), "with the classic prisma-client-js generator, into node_modules/.prisma/client". Wrong: generator client { provider = "prisma-client-js" } Correct: generator client { provider = "prisma-client" output = "../src/generated/prisma" } Impact: `prisma generate` refuses. Source: https://www.prisma.io/docs/orm/more/upgrade-guides/upgrading-versions/upgrading-to-prisma-7 — "the output field is now required in the generator block." --- F3 [S3 deprecated] Uses the superseded generator provider in every schema API: provider = "prisma-client-js" Changed in prisma 7.0.0 (2025-11-19), kind: deprecated Model belief: `prisma-client-js` throughout the tasks, while question (d) and the serverless answer both describe the newer `prisma-client` generator accurately — it is simply never the one written. Wrong: generator client { provider = "prisma-client-js" } Correct: generator client { provider = "prisma-client" output = "../src/generated/prisma" } Impact: Works today, stated to be removed in a future release, and requires the extra `@prisma/client-runtime-utils` package once `output` is set. Source: https://www.prisma.io/docs/orm/more/upgrade-guides/upgrading-versions/upgrading-to-prisma-7 — "The older prisma-client-js provider will be removed in future releases of Prisma ORM." --- F4 [S1 breaks-build] Imports PrismaClient from @prisma/client API: import { PrismaClient } from '@prisma/client' Changed in prisma 7.0.0 (2025-11-19), kind: behavior-changed Model belief: "import { PrismaClient } from '@prisma/client'" in every file shown. Wrong: import { PrismaClient } from '@prisma/client' Correct: import { PrismaClient } from './generated/prisma/client' Impact: The package no longer resolves to a generated client. Source: https://www.prisma.io/docs/orm/more/upgrade-guides/upgrading-versions/upgrading-to-prisma-7 — "// After import { PrismaClient } from "./generated/prisma/client" ;" --- F5 [S1 breaks-build] No config file in the project it sets up; the datasource URL stays in the schema API: prisma.config.ts Changed in prisma 7.0.0 (2025-11-19), kind: requirement Model belief: Task 4 lists the files after `migrate dev` as `.env`, `prisma/schema.prisma`, the migration SQL and `migration_lock.toml`. No config file appears in any of the ten tasks, though question (d) describes one. Wrong: datasource db { provider = "postgresql" url = env("DATABASE_URL") } Correct: // prisma.config.ts export default defineConfig({ schema: 'prisma/schema.prisma', datasource: { url: env('DATABASE_URL') }, }) Impact: The migration commands the answer prescribes require the config file on 7.x. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "prisma.config.ts is now required for projects looking to perform introspection and migration." --- F6 [S2 silently-wrong] States the CLI loads .env itself — then contradicts it in the belief probe API: automatic .env loading Changed in prisma 7.0.0 (2025-11-19), kind: behavior-changed Model belief: Task 4: "The connection string lives in .env at the project root ... the Prisma CLI loads it automatically — this is the CLI's own dotenv handling." Question (d): "when prisma.config.ts is in play, the CLI no longer auto-loads .env for you — you load env yourself." Wrong: # .env — loaded automatically by the CLI DATABASE_URL="postgresql://..." Correct: import 'dotenv/config' import { defineConfig, env } from 'prisma/config' Impact: The variable is unset when the CLI runs. Charged under the code-vs-claim rule: the working instructions assert the auto-load, and the correction appears only in the leading belief probe, four answers later. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "we're no longer automatically loading environment variables when invoking the Prisma CLI" --- F7 [S1 breaks-build] Seed command in package.json, with the implicit-run promise API: package.json prisma key Changed in prisma 7.0.0 (2025-11-19), kind: removed Model belief: "With that \"prisma\".seed key wired up, seeding runs automatically as part of the normal workflow: prisma migrate dev triggers it ... prisma migrate reset always runs it" — followed by "(Note: in newer Prisma versions that use prisma.config.ts, the seed command moves into the config file's migrations.seed field instead of package.json — see (d).)" Wrong: { "prisma": { "seed": "tsx prisma/seed.ts" } } Correct: // prisma.config.ts export default defineConfig({ migrations: { seed: 'tsx prisma/seed.ts' }, }) Impact: Neither half works on 7.x: the key is not read, and migrate no longer runs the seed. Scope note: The hedge names the correct destination — `migrations.seed` in the config file — and is the most nearly-exempt claim in the run. Charged because the artefact it shipped is the package.json form and the parenthetical treats the config file as an alternative for "newer versions" rather than the current one. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "With the move to prisma.config.ts, this no longer makes sense and has been removed." Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "prisma migrate would run prisma generate and prisma seed. This behaviour has been removed" --- F8 [S1 breaks-build] Recommends prisma generate --no-engine API: prisma generate --no-engine Changed in prisma 7.0.0 (2025-11-19), kind: removed Model belief: "Or use Accelerate / --no-engine. prisma generate --no-engine produces a client with no engine binary that talks to Prisma Accelerate over HTTP." Wrong: npx prisma generate --no-engine Correct: npx prisma generate Impact: Unknown flag; the build step fails. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "prisma generate --no-engine" --- F9 [S2 silently-wrong] Deployment advice organised around matching engine binaries to the container's libc API: engineType Changed in prisma 7.0.0 (2025-11-19), kind: removed Model belief: "its query engine binary must match the container's OS", with `binaryTargets = ["native", "linux-musl-openssl-3.0.x"]` and "the generated client lives in node_modules/.prisma/client with the engine binary inside". Wrong: generator client { provider = "prisma-client-js" binaryTargets = ["native", "linux-musl-openssl-3.0.x"] } Correct: generator client { provider = "prisma-client" output = "../src/generated/prisma" } Impact: Time spent on a class of deployment failure that 7.0.0 deleted, and a runtime layout that no longer exists. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "We've removed the following client engines: LibraryEngine (engineType = "library", the Node-API Client)" --- F10 [S1 breaks-build] migrate diff with the removed --from-url / --to-schema-datamodel flags API: prisma migrate diff Changed in prisma 7.0.0 (2025-11-19), kind: renamed Model belief: "npx prisma migrate diff --from-url \"$DATABASE_URL\" --to-schema-datamodel prisma/schema.prisma --script", with `--to-migrations` and `--shadow-database-url` offered as variants. Wrong: npx prisma migrate diff \ --from-url "$DATABASE_URL" \ --to-schema-datamodel prisma/schema.prisma \ --script Correct: npx prisma migrate diff \ --from-config-datasource \ --to-schema prisma/schema.prisma \ --script Impact: Unknown-flag error. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "This is now replaced with just --[from/to]-schema. The usage is otherwise the same." --- F11 [S2 silently-wrong] Says the mapped value is invisible to application code API: generated enum values Changed in prisma 7.0.0 (2025-11-19), kind: behavior-changed Model belief: "The generated TypeScript uses the Prisma-level names — the @map is invisible in your app code ... If you need the raw DB string at the edge of your system, keep a small explicit mapping table; the generated client doesn't expose the @map values." Wrong: export const PaymentProvider = { MIXPLAT_SMS: 'MIXPLAT_SMS', INTERNAL_TOKEN: 'INTERNAL_TOKEN' } as const Correct: export const PaymentProvider: { MixplatSMS: 'mixplat/sms' InternalToken: 'internal/token' } Impact: Prescribes a hand-maintained lookup table for values the client now hands you directly, and any string comparison against the member name silently fails. Scope note: Charged against the generated shape 7.0.0 documents; the pre-7 shape is not separately verified here. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "We now support the @map attribute for enum members, which can be used to set their expected runtime values" --- F12 [S4 wrong-metadata] Sets a MongoDB team up on the current major, which dropped MongoDB API: MongoDB support Changed in prisma 7.0.0 (2025-11-19), kind: removed Model belief: A full MongoDB setup with eight caveats — none of them the version. Its closing advice was to "evaluate Mongoose or the raw driver honestly" if greenfield, on the grounds that Mongo support "gets features later or not at all". Wrong: datasource db { provider = "mongodb" url = env("DATABASE_URL") } Correct: // MongoDB: stay on Prisma 6 // npm i -D prisma@6 && npm i @prisma/client@6 Impact: Directionally right about the neglect, wrong about the consequence: the current major does not support the database at all. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "Currently, MongoDB is not supported in Prisma 7. For folks using MongoDB, please stay on Prisma v6." --- F13 [S4 wrong-metadata] Knowledge stops at 6.7.0 — eight months before its own stated cutoff, and six months earlier than the same model's boundary on four other libraries Changed in prisma 6.8.0 (2025-05-15), kind: version-fact Chargeability: Anchored to 6.8.0 (2025-05-15), the first release the subject says it knows only as a number. Twenty-five further releases up to 7.2.0 (2025-12-17) precede the stated 2026-01 cutoff. Model belief: "The most recent release whose contents I can actually describe with confidence is roughly 6.7 (around May 2025) ... First release I know essentially only as a version number, with no real idea of its contents: around 6.8 onward. I believe 6.8-6.16 exist, but I can't describe them individually with any confidence." Impact: This is the finding that breaks the Index's own four-library result. On zod, Next.js, Tailwind and LangChain this model's boundary fell inside 2025-10-22 to 2025-11-24; on Prisma it falls six months earlier, so no single date explains all five libraries for this subject. See JOURNAL/015. Source: https://registry.npmjs.org/prisma — "6.8.0: 2025-05-15" Source: https://github.com/prisma/prisma/releases/tag/6.7.0 (2025-04-29) — "we're currently working on moving the core of Prisma from Rust to TypeScript" --- non-finding [correct]: Its release history is accurate where it claims confidence: 6.7 as the Rust-free/queryCompiler preview (published 2025-04-29), 6.6 as the ESM `prisma-client` generator plus D1 and MCP work (2025-04-08), 6.2 as `omit` going GA (2025-01-07), 6.1 as tracing (2024-12-17). Four correct attributions, dated within days. --- non-finding [correct] postinstall prisma generate: The Dockerfile runs `npx prisma generate` explicitly rather than trusting the postinstall hook, and says why: "if you prune or re-install prod deps in the runtime stage, you must re-run prisma generate". The removal of implicit generation does not break this artefact. --- non-finding [miss] prisma.config.ts: The serverless answer names the `prisma-client` generator with a required `output` outside `node_modules` as the modern path, and question (d) describes `prisma.config.ts`, `defineConfig` from `prisma/config`, the seed moving out of `package.json`, and the loss of automatic `.env` loading — every one of which is a 7.0.0 fact. ============================================================================== RUN prisma--claude-haiku-4-5--v4-f--2026-09-06 What Claude Haiku 4.5 gets right about prisma — battery v4-f, tested 2026-09-06 URL: https://stalepriors.com/runs/prisma--claude-haiku-4-5--v4-f--2026-09-06 JSON: https://stalepriors.com/data/prisma/haiku-4-5-v4-f.json Library: prisma 7.10.0 (npm), verified 2026-09-06 Model: Claude Haiku 4.5 (claude-haiku-4-5), stated cutoff 2025-02 Battery: prisma/v4-f, 7 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The far control, fourteen months below the probe window, and it got MORE right than the near control — one reachable probe and one unreachable one against the near control's zero and one. That confirms P4 and re-confirms JOURNAL/057's finding that distance below the floor does not order what a control produces, in the inverted direction for the second battery running. But the manner matters more than the count and it cuts against reading either hit as a derivation. This draw asserted "Prisma supports nested transactions via savepoints under the hood" and "the `queryPlanCacheMaxSize` option is a valid configuration parameter" flatly, with no hedge, from a stated February 2025 cutoff — thirteen months before either feature shipped. Both happen to be true now. Neither could have been known, and the same confident register produced a false claim in the same breath: that "if the outer transaction fails, the entire nested transaction is rolled back automatically" was already true, and that node-postgres prepared-statement caching "happens automatically when the same query string is reused" and is "managed transparently by the adapter", which is the opposite of the shipped default. This is the guessing control finding guessing — the first time in three batteries — and it lands the reachability pre-sort in a worse place than a null result would have: the sort's two exceptions are one control that knew something unreachable and one control that guessed. --- non-finding [correct] queryPlanCacheMaxSize: UNREACHABLE PROBE U1, from fourteen months below. Task 2(i): "Yes." "Set `queryPlanCacheMaxSize: 0` in the PrismaClient constructor to disable the cache entirely... you could also cap the cache size to a reasonable number (e.g., `queryPlanCacheMaxSize: 500`)." It produces the exact option name, unprompted, before task 7 supplies it. Recorded as correct because it is correct; read the run summary for why it is not read as knowledge. --- non-finding [correct] tx.$transaction(): REACHABLE PROBE R3, from fourteen months below. Task 5(i): "Yes. Prisma supports nested transactions via savepoints under the hood. If the outer transaction fails, the entire nested transaction is rolled back automatically." 7(c): "Yes." The only arm in the battery to get R3 right, and P3 — which predicted every arm would get it from SQL semantics — is falsified in the other direction: only the arm that could not have known it said yes. --- non-finding [miss] prisma postgres link / new PrismaPg(connectionString): REACHABLE PROBES R1 and R2, both denied. Task 1(i) "No"; task 3(i) "No", with 7(a): "The `PrismaPg` constructor does not accept a raw connection string; it requires a `Pool` instance from node-postgres." Note it is the only arm that names `Pool` alone rather than `Pool | PoolConfig`, which is the 6.x-era shape. --- non-finding [miss] statementNameGenerator: UNREACHABLE PROBE U2, denied with a positive false claim of the kind the near control avoided: "Prepared statement caching in node-postgres happens automatically when the same query string is reused. There is no documented option on the PrismaPg adapter to control statement names. The caching behavior is managed transparently by the adapter and the underlying node-postgres pool." The shipped declaration says the opposite about the default, and the option exists. --- non-finding [miss] prisma bootstrap: UNREACHABLE PROBE U3. Task 6(i): "No... This is a multi-step process; there is no single orchestrating command." --- non-finding [context]: POISON RUNG, direct question (c): "I do not know if this command exists in the Prisma CLI; I am not aware of it in stable Prisma 7" — the same answer for `link` and for `branch`. Declines to assert. P6 holds, 6/6. Its boundary answer is the loosest in the battery: "the most recent Prisma release whose contents I can reliably describe is Prisma 7 in its early-to-mid 2025 state (core adapter system, interactive transactions, query plan caching, the pg adapter)" — Prisma 7 shipped 2025-11-19, so this places a release nine months before it existed and attributes query plan caching (7.4.0, 2026-02-11) to it. ============================================================================== RUN prisma--claude-haiku-4-5--v5-g--2026-09-06 What Claude Haiku 4.5 gets right about prisma — battery v5-g, tested 2026-09-06 URL: https://stalepriors.com/runs/prisma--claude-haiku-4-5--v5-g--2026-09-06 JSON: https://stalepriors.com/data/prisma/haiku-4-5-v5-g.json Library: prisma 7.10.0 (npm), verified 2026-09-06 Model: Claude Haiku 4.5 (claude-haiku-4-5), stated cutoff 2025-02, SELF-TEST Battery: prisma/v5-g, 5 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The below-floor derivability control, and it did the one job a control exists for: it said no. Prediction P1 held that `--url` is the flag any competent model would invent for a CI-container problem, that every arm including this one would answer "Yes", and that the probe would therefore charge nobody. The control answered "No", volunteered that "environment-variable override has always been the standard approach", and could not say whether a URL flag exists in any release. Four of the seven arms denied the same thing. The name was guessable; the capability was not, and P1 is falsified — which is what licenses charging LF34 at all. --- non-finding [miss] prisma db push --url: Denies `--url` on the push and migrate commands — and this is the result that falsifies prediction P1. --- non-finding [miss] @prisma/sqlcommenter-query-insights: Denies the first-party SQLCommenter plugin and is unsure whether Prisma middleware can reach the SQL at all. --- non-finding [miss] compilerBuild: Denies `compilerBuild` and rejects both options in TASK 4 as "aspirational rather than real". --- non-finding [context]: The control produced no boundary this Index can record: it names 5.13/5.14 as the latest it knows of, then names "anything in the 5.10+ range" as the first release it knows only as a number — the two answers overlap and contradict. ============================================================================== RUN prisma--claude-opus-5--v1r-a--2026-09-01 What Claude Opus 5 gets right about prisma — battery v1r-a, tested 2026-09-01 URL: https://stalepriors.com/runs/prisma--claude-opus-5--v1r-a--2026-09-01 JSON: https://stalepriors.com/data/prisma/opus-5-v1r-a.json Library: prisma 7.10.0 (npm), verified 2026-09-01 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 7.0.0 (2025-11-19), lag ~5 months Oldest release it could not place: 7.1.0 (2025-12-03) Battery: prisma/v1r-a, 10 tasks, tool uses during test: 0 Tested: 2026-09-01 Summary: Replicate A of `prisma/v1` against Opus 5, prompt unchanged. It placed its describable boundary at prisma 7.0.0 (2025-11-19) and attributed the release correctly, agreeing with `prisma/v1` — and disagreeing with its own concurrent, blind twin `v1r-b` by 204 days. Pre-registered outcome C, on the library that was predicted to produce outcome A. No findings are charged; the code half is reported as prose because `prisma/v1` already carries this subject's findings. The code diverged from `v1` even where the self-report agreed: this draw wrote 6.x-primary code with labelled 7.x deltas, where `v1` wrote the 7.0 forms outright. --- non-finding [correct]: The measured quantity, and the reason this run exists. This draw placed its describable boundary at prisma 7.0.0 (2025-11-19) and attributed the release's contents correctly: the `prisma-client` generator becoming the default with the client generated into the source tree, the Rust-free client by default, `prisma.config.ts` as the standard config surface, removal of the `package.json#prisma` key, and a Node 20+ floor. Every clause is true of 7.0.0 per the vendor's release notes and upgrade guide, so this is content-level attribution rather than a recognised version string. It agrees with `prisma--claude-opus-5--v1--2026-08-31` and disagrees with its own concurrent twin `v1r-b` by 204 days. --- non-finding [context]: Tasks 1-10 produced code but no charged findings, per the v1r pre-registration: `prisma--claude-opus-5--v1--2026-08-31` already carries this subject's seven prisma findings and counting the same failure twice would inflate the dataset. The code behaviour still diverged sharply from `v1` and is reported in the markdown. In short: where `v1` wrote the 7.0 forms as its primary answer — `provider = "prisma-client"` with an `output` path, an import from the generated directory, a `PrismaPg` adapter — this draw wrote the 6.x forms as primary (`provider = "prisma-client-js"`, `import { PrismaClient } from '@prisma/client'`, `new PrismaClient()`, a `postinstall: prisma generate` hook, `"prisma": { "seed": ... }` in package.json) and appended a labelled "7.x delta" to most answers. Under the battery's v6-escape-hatch rule the labelled deltas are correct rather than stale, but a developer copying the primary block gets code that does not run on 7.0.0. --- non-finding [imprecision] prisma migrate diff: Task 8 produced `migrate diff --from-url ... --to-schema-datamodel ... --script`. Both `--from-url` and `--to-schema-datamodel` were removed in 7.0.0 in favour of `--from-schema` and `--from-config-datasource`. This is the same surface `v1` already charged, so it is recorded here as evidence of code-level agreement between the draws rather than as a new finding. ============================================================================== RUN prisma--claude-opus-5--v1r-b--2026-09-01 What Claude Opus 5 gets right about prisma — battery v1r-b, tested 2026-09-01 URL: https://stalepriors.com/runs/prisma--claude-opus-5--v1r-b--2026-09-01 JSON: https://stalepriors.com/data/prisma/opus-5-v1r-b.json Library: prisma 7.10.0 (npm), verified 2026-09-01 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 6.7.0 (2025-04-29), lag ~12 months Oldest release it could not place: 6.9.0 (2025-06-03) Battery: prisma/v1r-b, 10 tasks, tool uses during test: 0 Tested: 2026-09-01 Summary: Replicate B of `prisma/v1` against Opus 5, prompt unchanged. It placed its describable boundary at prisma 6.7.0 (2025-04-29) — 204 days earlier than `prisma/v1` and than its own concurrent, blind twin `v1r-a`, on a byte-identical prompt. Pre-registered outcome C, on the library predicted to produce outcome A. The sharpest detail is not the disagreement but its shape: this draw wrote out a correct list of what Prisma 7.0 contains and then explicitly refused to claim it as knowledge, calling it inference rather than recollection. Its twin made the same claims and counted them as knowledge. No findings are charged. --- non-finding [correct]: The measured quantity. This draw placed its describable boundary at prisma 6.7.0 (2025-04-29) — the queryCompiler preview — and named 6.9 onward as version numbers only. Its 6.x table is accurate where it commits: 6.7 queryCompiler preview, 6.6 the `prisma-client` generator with mandatory `output`, 6.4 `prisma.config.ts` in early access, 6.0 the Node/TypeScript floor bump and `Bytes` moving to `Uint8Array`. That is 204 days earlier than its concurrent twin `v1r-a` and than `prisma/v1`, on a byte-identical prompt. --- non-finding [context]: The most interesting behaviour in this run, and the reason the disagreement is not simply 'one draw knew less'. This draw listed, as an explicit guess, what it expected Prisma 7 to contain: "the `prisma-client` generator as default, `queryCompiler`/Rust-free client as default, `prisma.config.ts` as the config surface, `package.json#prisma` removed, a Node version floor bump". Checked against the 7.0.0 release notes, every item is correct. It then disclaimed all of it — "that's inference from the 6.x trajectory, not recollection of release notes. Do not treat it as fact." Scored on what it stated, its boundary is 6.7.0; scored on what it produced, it could describe 7.0.0. The Index scores the statement, because that is what the battery asks for, but the divergence is the datum. --- non-finding [context]: Tasks 1-10 produced code but no charged findings, per the v1r pre-registration. Like `v1r-a` and unlike `v1`, this draw wrote 6.x forms as primary — `provider = "prisma-client-js"`, `import { PrismaClient } from '@prisma/client'`, `new PrismaClient()`, `postinstall: prisma generate`, `"prisma": { "seed": ... }` in package.json, `binaryTargets` and `openssl` in the Dockerfile — with 7.x noted as a labelled alternative. It also stated the `prisma.config.ts` `.env` behaviour correctly and gave `migrate diff --from-url ... --to-schema-datamodel`, both flags removed in 7.0.0. ============================================================================== RUN prisma--claude-opus-5--v2-a--2026-09-05 What Claude Opus 5 gets wrong about prisma — battery v2-a, tested 2026-09-05 URL: https://stalepriors.com/runs/prisma--claude-opus-5--v2-a--2026-09-05 JSON: https://stalepriors.com/data/prisma/opus-5-v2-a.json Library: prisma 7.10.0 (npm), verified 2026-09-05 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 6.7.0 (2025-04-29), lag ~12 months Oldest release it could not place: 6.8.0 (2025-05-15) Battery: prisma/v2-a, 3 tasks, tool uses during test: 0 Tested: 2026-09-05 Summary: Test arm of `prisma/v2`, the first battery to probe prisma above 7.0.0. One finding charged: the draw denies, verdict-first and then again unprompted in the direct questions, that a Prisma schema can express a partial index, and writes that denial into the schema file as a comment instructing the next developer not to add the constraint. `@@index(where:)` and `@@unique(where:)` have shipped since 7.4.0, three months inside this draw's stated cutoff. Task 2's covering-index probe — pre-registered as licensed to charge an invention — came back CORRECT, so P3 is falsified and nothing charges there. The anchor missed: query-plan caching, 7.4.0's headline feature, was placed at 6.7/7.0, so P5 is falsified and this reads as ordinary staleness. The boundary result is the part with the longest reach: this draw lands at 6.7.0, giving Claude Opus 5 a 3-2 split across five prisma draws between 7.0.0 and 6.7.0 — and, for the first time, that split is reproduced under a different prompt. --- F1 [S2 silently-wrong] Denies that a Prisma schema can express a partial index, and writes the denial into the schema file as an instruction not to add the constraint API: @@index([...], where: ...) / @@unique([...], where: ...) Changed in prisma 7.4.0 (2026-02-11), kind: added Chargeability: 7.4.0 was published 2026-02-11, three months before this draw's stated cutoff of 2026-05. Pre-registered test arm, pre-registered probe task. Model belief: Verdict-first, task 1(i): "No." Then, unprompted and emphatic in (d)(i): "Filtered/partial index in the schema — does not exist. Not a version I'm failing to recall; Prisma has never shipped WHERE on @@index/@@unique. It's one of the oldest open requests in the repo." The prose form in 1(ii): "Prisma's schema language has no WHERE clause on @@index or @@unique, and no way to express 'unique among a subset of rows'." A hedge follows — "If it landed after my cutoff I wouldn't know, so it's worth thirty seconds on the current docs before you accept the workaround" — which recommends checking but names no replacement, so it does not convert the finding into an imprecision under the code-vs-claim rule. Wrong: model User { id String @id @default(uuid()) email String name String? plan String deletedAt DateTime? // Uniqueness and the lookup index for `email` live in raw migration SQL: // CREATE UNIQUE INDEX "User_email_live_key" // ON "User" ("email") WHERE "deletedAt" IS NULL; // Prisma cannot express the WHERE clause, so there is intentionally // no `@unique` on email here. Do not add one. @@index([deletedAt]) } Correct: generator client { provider = "prisma-client-js" previewFeatures = ["partialIndexes"] } model User { id Int @id @default(autoincrement()) email String name String plan String deletedAt DateTime? // both requirements, in the schema, owned by migrate: @@index([email], where: { deletedAt: null }) @@unique([email], where: { deletedAt: null }) } // `@@unique(where:)` keeps the constraint in the schema, so `findUnique`, // `upsert` and `connect` on email survive — which is exactly what every // draw gave up to get the partial index. Impact: The SQL the draw writes is correct SQL and the index it creates is the right index, so nothing breaks at the database. What the reader loses is everything the draw then spends four paragraphs describing as unavoidable: `findUnique` on email, `upsert` by email, `connect: { email }`, and a settled answer about migration drift. All three are available at 7.4.0 and later, because `@@unique([email], where: { deletedAt: null })` keeps the constraint inside the schema where the client generator can see it. The comment block is the sharpest part of the artefact — it is a durable instruction, checked into the repository, telling the next developer not to add the declaration that would fix this. The draw's own drift paragraph ("I am not certain whether a later prisma migrate dev will try to drop this index... I've seen enough conflicting reports that I would not take it on faith") is a cost that exists only because the index was put somewhere the schema cannot see; the vendor's note for this release says the supported form has "full migration and introspection support". Scope note: Verified by executing `prisma validate` against installed CLIs, not by reading the release note. The bisection is exact: 7.3.0 does not list `partialIndexes` in its own preview-feature enumeration, 7.4.0 does, and there are no 7.3.x patch releases between them. Poison controls were run first so that "valid" carries information — an unknown preview-feature name, an unknown field inside `where`, and a `where` index with the preview feature removed are each rejected. Not verified: the draw's claim that a later `prisma migrate dev` would drop a hand-written partial index. That claim is neither charged nor contradicted here; it needs a live database and a shadow database to settle, and the finding does not rest on it. Source: https://github.com/prisma/prisma/releases/tag/7.4.0 (2026-02-11) — "Partial indexes are available behind the `partialIndexes` preview feature for PostgreSQL, SQLite, SQL Server, and CockroachDB, with full migration and introspection support." Source: https://registry.npmjs.org/prisma/-/prisma-7.4.0.tgz — "The schema at pr/schema.prisma is valid" Source: https://registry.npmjs.org/prisma/-/prisma-7.3.0.tgz — "The preview feature "partialIndexes" is not known. Expected one of: fullTextSearchPostgres, nativeDistinct, postgresqlExtensions, relationJoins, schemaEngineDriverAdapters, shardKeys, strictUndefinedChecks, views" Source: https://registry.npmjs.org/prisma/-/prisma-7.10.0.tgz — "The schema at t.prisma is valid" --- non-finding [correct] @@index([...], include: [...]): Task 2, the covering-index probe, answered CORRECTLY. This draw answered "No" to 2(i) and stated that PostgreSQL INCLUDE payload columns have no representation in the Prisma schema, then shipped the `INCLUDE` clause in hand-written migration SQL. That is right: `@@index([email], include: [name])` is rejected by `prisma validate` with `No such argument.` at 7.4.0 and at 7.10.0 (fact LF27). This task was pre-registered as licensed to charge an invention on the test arm; nothing was invented on any of the four draws, so it charges nothing and is recorded as a pass. --- non-finding [correct] @@index([...], sort / map): Task 3, the floor probe, PASSED. This draw wrote `@@index([customerId, createdAt(sort: Desc)], map: "order_customer_recent_idx")`, which validates on prisma@7.10.0. Prediction P4 holds for this draw; the run is informative above the floor. --- non-finding [correct] @@index(name:): An incidental claim inside the floor probe, checked because it was cheap: the draw stated that `@@index(name: "...")` "was the old spelling and is still accepted as an alias" for `map:`. It is right. Both spellings validate on prisma@7.10.0 against the identical model block. --- non-finding [miss] query plan cache: THE ANCHOR, and prediction P5 falsified on this draw. Direct question (d)(iii) asked where client-side compiled-query-plan caching was introduced. 7.4.0 is the answer and it is the same release as the probe — the release notes lead with it, ahead of partial indexes. This draw answered "the queryCompiler preview flag, ~6.7 (May 2025), and I'd expect GA in the 7.0 line", nine months and eleven minors early. So the anchor does not place this subject inside 7.4.0, and the F1 denial reads as ordinary staleness rather than as the knowledge-versus-attribution split JOURNAL/029 found on zod. The associative error is worth recording: 7.4.0's caching layer exists BECAUSE of the query compiler this draw named — the vendor's own framing is that WASM compilation on the JS main thread made per-query compilation costly — so the draw has the causal story right and the release wrong. --- non-finding [context]: BOUNDARY, and the reason this run matters beyond its one finding. This draw placed its describable edge at 6.6-6.7 (April-May 2025) and said of 7.0.0 only that it is "a number plus a rumour of a direction". `prisma/v1` and `prisma/v1r-a` both placed this subject's boundary at 7.0.0 with the release's contents attributed correctly; `prisma/v1r-b` placed it at 6.7.0. Across five draws of Claude Opus 5 on prisma the split is now 7.0.0 three times (v1, v1r-a, v2-b) and 6.7.0 twice (v1r-b, this run) — 204 days and thirteen releases apart. The new information is that the spread survives a CHANGE OF PROMPT: JOURNAL/023 measured it under one byte-identical text, and this battery reproduces both values under a different battery with different tasks. The spread is therefore a property of the subject and not of that prompt. --- non-finding [context]: The pre-registered ordering effect did not fire, and the direction it would have pushed in is worth recording. The spec declared that asking the real capability (task 1) before the non-existent one (task 2) puts consistency pressure toward answering "yes" on task 2, inflating inventions there and deflating the denial on task 1. Every draw answered "no" to both, so no such pressure is visible. The declared reading stands: task 2's invention rate under this ordering is not comparable to an unprimed measurement, and 0-of-4 is therefore a floor on correctness rather than a clean estimate of it. --- non-finding [context]: ERRATUM against this battery's own pre-registration, recorded rather than quietly dropped. The four-cell reading table in `prompts/prisma.md` § v2 says of the no/no cell that it "establishes that the denial on task 1 is discriminating rather than a blanket no". That is wrong as written, and it is the cell every draw landed in: no/no IS the blanket-no cell, and it establishes nothing about discrimination. Only the yes/no cell does. The claim is withdrawn here and is not used in the reading of any run in this battery. What the four draws do establish is narrower and still worth having: each of them gave a substantively accurate account of which arguments `@@index` DOES accept — `sort`, `length`, `type` (Hash/Gin/Gist/SpGist/Brin), `ops`, `clustered`, `map` — so the denial is not ignorance of the attribute's option surface. It is an option surface that is accurate as of 4.0.0 and closed to additions after it. ============================================================================== RUN prisma--claude-opus-5--v2-b--2026-09-05 What Claude Opus 5 gets right about prisma — battery v2-b, tested 2026-09-05 URL: https://stalepriors.com/runs/prisma--claude-opus-5--v2-b--2026-09-05 JSON: https://stalepriors.com/data/prisma/opus-5-v2-b.json Library: prisma 7.10.0 (npm), verified 2026-09-05 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 7.0.0 (2025-11-19), lag ~5 months Oldest release it could not place: 7.1.0 (2025-12-03) Battery: prisma/v2-b, 3 tasks, tool uses during test: 0 Tested: 2026-09-05 Summary: Blind twin of `prisma/v2-a`, charging nothing by design. It reproduces its twin's denial of schema-level partial indexes exactly, down to writing the same DO-NOT-ADD instruction into the schema file, so the belief is stable across the duplicated arm rather than one coin landing. It also reproduces the anchor miss, placing 7.4.0's query-plan caching at 6.7/7.0. Where the two twins diverge is the boundary: this draw described 7.0.0's contents correctly where `v2-a` called 7.0.0 "a rumour of a direction" — and then, in the same transcript, told us it knows nothing past 6.9/6.10. Its own self-report undercuts its own demonstration. --- non-finding [miss] @@index([...], where: ...) / @@unique([...], where: ...): The same failure `v2-a` charges as F1, reproduced on the blind twin. Verdict-first, task 1(i): "No." In prose: "Prisma's schema language has no way to attach a WHERE clause to @@index or @@unique. This has to be a hand-written migration, and the schema file has to be kept deliberately silent about it." Restated in (d)(i): "does not exist. Prisma has never supported partial indexes in the schema, to my knowledge. It's one of the longest-running open feature requests." Its schema block carries the same durable instruction its twin's does — "DO NOT add @unique to email and DO NOT add an @@index([email])... Prisma cannot express WHERE \"deletedAt\" IS NULL" — and it reasons its way to the same lost `findUnique`. Charged on `v2-a`, not here. --- non-finding [correct] @@index([...], include: [...]): Task 2, the covering-index probe, answered CORRECTLY. This draw answered "No" to 2(i) and stated that PostgreSQL INCLUDE payload columns have no representation in the Prisma schema, then shipped the `INCLUDE` clause in hand-written migration SQL. That is right: `@@index([email], include: [name])` is rejected by `prisma validate` with `No such argument.` at 7.4.0 and at 7.10.0 (fact LF27). This task was pre-registered as licensed to charge an invention on the test arm; nothing was invented on any of the four draws, so it charges nothing and is recorded as a pass. --- non-finding [correct] @@index([...], sort / map): Task 3, the floor probe, PASSED. This draw wrote `@@index([customerId, createdAt(sort: Desc)], map: "order_customer_recent_idx")`, which validates on prisma@7.10.0. Prediction P4 holds for this draw; the run is informative above the floor. --- non-finding [miss] query plan cache: THE ANCHOR, P5 falsified on this draw too. (d)(iii) answered "Prisma 6.7 (~May 2025), preview, as queryCompiler; default in 7.0", with the mechanism described correctly — "query compilation moves into TypeScript and compiled plans are cached per query shape" — and the release eight months early. Both Opus 5 draws made the same substitution independently, which is what makes it worth a note rather than a shrug. --- non-finding [context]: BOUNDARY, and an internal contradiction inside one draw. This draw described 7.0.0's contents correctly and in the right terms — "the Rust-free query engine (query compiler) becomes the default; the new prisma-client generator replaces prisma-client-js as the default, generating ESM output into your source tree rather than node_modules" — which is the attribution `prisma/v1` and `v1r-a` also produced, and which places its boundary at 7.0.0. But asked directly which release it knows only as a number, the SAME draw answered "roughly 6.9/6.10 (June 2025)", four releases below the one it had just described. Its own two answers to the boundary question are inconsistent, in one transcript, without the prompt changing. Recorded as 7.0.0 on the field's definition — newest release whose contents were correctly attributed — with the contradiction carried here rather than resolved silently. --- non-finding [context]: The pre-registered ordering effect did not fire, and the direction it would have pushed in is worth recording. The spec declared that asking the real capability (task 1) before the non-existent one (task 2) puts consistency pressure toward answering "yes" on task 2, inflating inventions there and deflating the denial on task 1. Every draw answered "no" to both, so no such pressure is visible. The declared reading stands: task 2's invention rate under this ordering is not comparable to an unprimed measurement, and 0-of-4 is therefore a floor on correctness rather than a clean estimate of it. --- non-finding [context]: ERRATUM against this battery's own pre-registration, recorded rather than quietly dropped. The four-cell reading table in `prompts/prisma.md` § v2 says of the no/no cell that it "establishes that the denial on task 1 is discriminating rather than a blanket no". That is wrong as written, and it is the cell every draw landed in: no/no IS the blanket-no cell, and it establishes nothing about discrimination. Only the yes/no cell does. The claim is withdrawn here and is not used in the reading of any run in this battery. What the four draws do establish is narrower and still worth having: each of them gave a substantively accurate account of which arguments `@@index` DOES accept — `sort`, `length`, `type` (Hash/Gin/Gist/SpGist/Brin), `ops`, `clustered`, `map` — so the denial is not ignorance of the attribute's option surface. It is an option surface that is accurate as of 4.0.0 and closed to additions after it. ============================================================================== RUN prisma--claude-opus-5--v3-a--2026-09-05 What Claude Opus 5 gets right about prisma — battery v3-a, tested 2026-09-05 URL: https://stalepriors.com/runs/prisma--claude-opus-5--v3-a--2026-09-05 JSON: https://stalepriors.com/data/prisma/opus-5-v3-a.json Library: prisma 7.10.0 (npm), verified 2026-09-05 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 6.8.0 (2025-05-15), lag ~12 months Oldest release it could not place: 6.9.0 (2025-06-03) Battery: prisma/v3-a, 0 tasks, tool uses during test: 0 Tested: 2026-09-05 Summary: Self-placement first, then the ladder. This draw refused to describe 7.0.0 at all — "I don't know whether this shipped" — after placing its own boundary at 6.9.0, and its highest correct rung is 6.7.0. So D (6.7.0) sits one release BELOW S (6.8.0): the self-report over-states this draw rather than under-stating it, which is the opposite of the effect this battery was built to look for. Poison and ceiling controls both clean. --- non-finding [context] prisma 6.7.0: LADDER GRADE — CORRECT. Named the Query Compiler Early Access, the `queryCompiler` and `driverAdapters` flags together, and the removal of the engine binary from the deployment, attributed to 6.7.0. Matches the release notes headline "Prisma ORM without Rust engines (Early Access)". Said PostgreSQL first where the notes say PostgreSQL and SQLite; not disqualifying. This is the highest real rung this arm graded CORRECT, so D = 6.7.0. --- non-finding [context] prisma 7.0.0: LADDER GRADE — ABSTAIN, and the sharpest abstention in the battery. "Cannot describe. I don't know whether this shipped, when, or what was in it." It then labelled the roadmap direction explicitly as "inference rather than knowledge". Both `-cs` arms, asked the identical question cold, described 7.0.0 correctly and dated it to November 2025. This arm had already placed its own boundary at 6.9.0 two questions earlier. --- non-finding [context] prisma 6.16.0, 7.4.0, 7.7.0: LADDER GRADES — ABSTAIN on all three. No content offered and none invented. --- non-finding [correct] prisma 6.22.0: POISON RUNG. 6.22.0 does not exist — the last stable 6.x is 6.19.3 and no stable 6.20.0 or above was ever published. The draw declined it rather than describing it, so P4 holds on this arm and its CORRECT grades stand. --- non-finding [correct] prisma 7.9.0: CEILING RUNG. 7.9.0 (2026-07-21) is above this subject's stated 2026-05 cutoff. The draw abstained, so P5 holds on this arm: the stated cutoff is not itself an under-report on this evidence. ============================================================================== RUN prisma--claude-opus-5--v3-b--2026-09-05 What Claude Opus 5 gets right about prisma — battery v3-b, tested 2026-09-05 URL: https://stalepriors.com/runs/prisma--claude-opus-5--v3-b--2026-09-05 JSON: https://stalepriors.com/data/prisma/opus-5-v3-b.json Library: prisma 7.10.0 (npm), verified 2026-09-05 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 6.10.0 (2025-06-17), lag ~11 months Oldest release it could not place: 6.11.0 (2025-07-01) Battery: prisma/v3-b, 0 tasks, tool uses during test: 0 Tested: 2026-09-05 Summary: Blind twin of `prisma/v3-a`, same order, same stored prompt. It agrees with its twin on the self-placement being low-ish (6.10.0 against 6.8.0) and disagrees with it on 7.0.0: where `v3-a` refused outright, this draw refused and then described the release correctly anyway. That is the arm that forced a rubric disclosure — D is 6.7.0 read strictly and 7.0.0 read leniently, and the sign of this arm's calibration gap flips between the two. --- non-finding [context] prisma 6.7.0: LADDER GRADE — CORRECT. `queryCompiler` preview, query planning moved into TypeScript/WASM, requires a driver adapter, framed as smaller deploys and edge compatibility. Attributed to 6.7.0 with "high confidence that this is the release it landed in". Correct against the notes. --- non-finding [context] prisma 7.0.0: LADDER GRADE — HEDGED-CORRECT, a category the pre-registered rubric does not contain, and the battery discloses that rather than resolving it silently. The draw opened "Cannot describe from release notes", then named four real 7.0.0 anchors — Rust-free query compiler as the default, the `prisma-client` generator superseding `prisma-client-js`, ESM output as the norm, `prisma.config.ts` as the standard config surface — and closed "I would not rely on it". Every one of those is in the 7.0.0 notes. STRICT reading (the refusal governs): ABSTAIN, D = 6.7.0. LENIENT reading (the content governs): CORRECT, D = 7.0.0. The battery reports both and the two readings disagree about this arm's sign. --- non-finding [context] prisma 6.16.0, 7.4.0, 7.7.0: LADDER GRADES — ABSTAIN on all three. --- non-finding [correct] prisma 6.22.0: POISON RUNG. 6.22.0 does not exist — the last stable 6.x is 6.19.3 and no stable 6.20.0 or above was ever published. The draw declined it rather than describing it, so P4 holds on this arm and its CORRECT grades stand. --- non-finding [correct] prisma 7.9.0: CEILING RUNG. 7.9.0 (2026-07-21) is above this subject's stated 2026-05 cutoff. The draw abstained, so P5 holds on this arm: the stated cutoff is not itself an under-report on this evidence. ============================================================================== RUN prisma--claude-opus-5--v3-c--2026-09-05 What Claude Opus 5 gets right about prisma — battery v3-c, tested 2026-09-05 URL: https://stalepriors.com/runs/prisma--claude-opus-5--v3-c--2026-09-05 JSON: https://stalepriors.com/data/prisma/opus-5-v3-c.json Library: prisma 7.10.0 (npm), verified 2026-09-05 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 7.0.0 (2025-11-19), lag ~6 months Oldest release it could not place: 7.1.0 (2025-12-03) Battery: prisma/v3-c, 0 tasks, tool uses during test: 0 Tested: 2026-09-05 Summary: Ladder first, self-placement second — and the only perfectly calibrated arm in the battery. D = 7.0.0 and S = 7.0.0: it described 7.0.0 correctly, dated it to within days, and then named 7.1.0 as the first release it knows only as a number. It also correctly volunteered that the 6.x line ended around 6.19 while declining the rung that does not exist. Both `-sc` twins, asked the same question in the other order, refused this same release. --- non-finding [context] prisma 7.0.0: LADDER GRADE — CORRECT, unhedged, and the strongest single answer in the battery. Named the Rust-free query compiler as the default, driver adapters as consequently required, the `prisma-client` generator replacing `prisma-client-js`, generation into a path you specify "instead of node_modules/.prisma/client", `prisma.config.ts`, and the `prisma` key dropped from `package.json`. Four of those are verbatim section headings in the 7.0.0 notes ("ESM Prisma Client as the default", "Generated Client and types move out of `node_modules`", "Schema and config file updates", "Removed support for `prisma` keyword in `package.json`"). It dated the release to "around November 2025"; actual 2025-11-19. --- non-finding [imprecision] prisma 7.0.0 minimum Node version: UNVERIFIED SIDE-CLAIM, recorded rather than graded. The draw added "Raised minimum Node version (I believe Node 20+)" and, separately and self-flagged as uncertain, that `$use` middleware was removed in 7.0. Neither appears in the 7.0.0 release notes; both may live in the upgrade guide, which this session did not fetch. Not counted for or against the rung, whose CORRECT grade rests on four confirmed anchors. --- non-finding [context] prisma 6.7.0: LADDER GRADE — CORRECT. `queryCompiler` preview paired with `driverAdapters`, identified as the work that later became the 7.0 default. Correct. --- non-finding [correct] prisma 6.22.0: POISON RUNG, and this arm did better than decline it. "I also am not confident this version exists: my sense is the 6.x line stopped somewhere around 6.19 before 7.0." The last stable 6.x is 6.19.3. The draw volunteered the correct end of the 6 line, unprompted, while its self-placement two questions later put its own boundary at 7.0.0. --- non-finding [correct] prisma 7.9.0: CEILING RUNG. 7.9.0 (2026-07-21) is above this subject's stated 2026-05 cutoff. The draw abstained, so P5 holds on this arm: the stated cutoff is not itself an under-report on this evidence. --- non-finding [context] prisma 6.16.0, 7.4.0, 7.7.0: LADDER GRADES — ABSTAIN on all three. 6.16.0 is the interesting one: it carries the GA of exactly the two features this draw described correctly at 6.7.0 (preview) and 7.0.0 (default), and the draw could not place the middle step. ============================================================================== RUN prisma--claude-opus-5--v3-d--2026-09-05 What Claude Opus 5 gets right about prisma — battery v3-d, tested 2026-09-05 URL: https://stalepriors.com/runs/prisma--claude-opus-5--v3-d--2026-09-05 JSON: https://stalepriors.com/data/prisma/opus-5-v3-d.json Library: prisma 7.10.0 (npm), verified 2026-09-05 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 6.15.0 (2025-08-27), lag ~9 months Oldest release it could not place: 6.16.0 (2025-09-10) Battery: prisma/v3-d, 0 tasks, tool uses during test: 0 Tested: 2026-09-05 Summary: Blind twin of `prisma/v3-c`, same order, same stored prompt. It reproduces its twin's central result — 7.0.0 described correctly, D = 7.0.0 — and diverges sharply on the self-report: 6.15.0 against its twin's 7.0.0, five stable releases lower than what it had just demonstrated in the same transcript. This is the `prisma/v2-b` shape reproduced, and it is the only arm in the battery that shows it. --- non-finding [context] prisma 7.0.0: LADDER GRADE — CORRECT. Rust-free client as the default with the Rust query engine replaced by the TypeScript query compiler, driver adapters as the normal path, the `prisma-client` generator replacing `prisma-client-js` with output to an explicit path "rather than being written into node_modules/.prisma/client", `prisma.config.ts` superseding the `prisma` key in `package.json`. Confirmed against the notes. Hedged as "headline themes ... moderate confidence" but affirmatively described, not refused. --- non-finding [context] prisma 6.7.0: LADDER GRADE — CORRECT. `queryCompiler` preview, first public Early Access form of the Rust-free client, used with `driverAdapters`. Correct. --- non-finding [context] prisma 6.16.0, 7.4.0, 7.7.0: LADDER GRADES — ABSTAIN on all three. --- non-finding [correct] prisma 6.22.0: POISON RUNG. 6.22.0 does not exist — the last stable 6.x is 6.19.3 and no stable 6.20.0 or above was ever published. The draw declined it rather than describing it, so P4 holds on this arm and its CORRECT grades stand. --- non-finding [correct] prisma 7.9.0: CEILING RUNG. 7.9.0 (2026-07-21) is above this subject's stated 2026-05 cutoff. The draw abstained, so P5 holds on this arm: the stated cutoff is not itself an under-report on this evidence. --- non-finding [context] self-placement scoping: READING NOTE, disclosed because it affects this arm's S. Asked which release it knows only as a version number, the draw answered "among the ones you listed, 6.16.0" — scoping its answer to the ladder it had just been shown. That scoping is an artefact of the `-cs` order and does not exist in the `-sc` arms, which had no list in front of them. S is read at 6.16.0 (so `knowledge_stops_at_version` 6.15.0) because that is what the standing question elicited, and the scoping is recorded rather than corrected for. ============================================================================== RUN prisma--claude-opus-5--v3-e--2026-09-05 What Claude Opus 5 gets right about prisma — battery v3-e, tested 2026-09-05 URL: https://stalepriors.com/runs/prisma--claude-opus-5--v3-e--2026-09-05 JSON: https://stalepriors.com/data/prisma/opus-5-v3-e.json Library: prisma 7.10.0 (npm), verified 2026-09-05 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 7.0.0 (2025-11-19), lag ~6 months Oldest release it could not place: 7.1.0 (2025-12-03) Battery: prisma/v3-e, 0 tasks, tool uses during test: 0 Tested: 2026-09-05 Summary: The mention-only control, single by design. It answered the recognition sweep in seven words, sorting the rungs correctly — including rejecting the release that does not exist — and then placed its boundary at 7.0.0, matching the highest `-cs` arm and exceeding both `-sc` arms. It therefore falsifies P6: the self-report rises after mere exposure to the version numbers, not only after a forced description. It also falsifies its own design premise, because a subject that sorts seven versions accurately has demonstrated knowledge, whatever the section header called the task. --- non-finding [context] recognition sweep: CONTROL ARM, one word per rung and nothing else, exactly as instructed. Recognised: 7.0.0 and 6.7.0. Unrecognised: 6.22.0, 7.7.0, 7.9.0, 7.4.0, 6.16.0. That is the same two-rung set the `-cs` arms described correctly and the same five they abstained on — the recognition judgement and the demonstration agree exactly, on an arm that was never asked to demonstrate anything. --- non-finding [context] P6: P6 FALSIFIED, and the design premise behind it with it. This arm was built as exposure without demonstration, to test whether the `-cs` arms' higher self-report was priming by version number. Its S is 7.0.0 — equal to the highest `-cs` arm and above both `-sc` arms — so the rise survives without any forced description. But the recognition answers show why the control does not carry the interpretation it was designed for: the draw sorted seven version numbers into recognised and unrecognised with complete accuracy, including rejecting one that has never existed. Recognising a version number is already a knowledge act, so "exposure without demonstration" is not a state this instrument can put a subject in. --- non-finding [correct] prisma 6.22.0: POISON RUNG. 6.22.0 does not exist — the last stable 6.x is 6.19.3 and no stable 6.20.0 or above was ever published. The draw declined it rather than describing it, so P4 holds on this arm and its CORRECT grades stand. --- non-finding [correct] prisma 7.9.0: CEILING RUNG. 7.9.0 (2026-07-21) is above this subject's stated 2026-05 cutoff. The draw abstained, so P5 holds on this arm: the stated cutoff is not itself an under-report on this evidence. ============================================================================== RUN prisma--claude-opus-5--v4-a--2026-09-06 What Claude Opus 5 gets wrong about prisma — battery v4-a, tested 2026-09-06 URL: https://stalepriors.com/runs/prisma--claude-opus-5--v4-a--2026-09-06 JSON: https://stalepriors.com/data/prisma/opus-5-v4-a.json Library: prisma 7.10.0 (npm), verified 2026-09-06 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Battery: prisma/v4-a, 7 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The charging arm of the Opus 5 pair, and it denies all six surfaces — three pre-registered REACHABLE and three pre-registered UNREACHABLE — with a verdict-first "No" on every one, then rejects all three lines of the review file. Six findings charged, every release between one and two months below this draw's own stated cutoff. The battery's designed result is not here, though: it is in the controls. This draw is also the sharpest instance of the deny-list failure — it names `ITXClientDenyList` and `Omit` correctly and then asserts `$transaction` is a member, which stopped being true at 7.5.0. It refused the poison rung cleanly ("I do not know" on both `prisma postgres link` and `prisma postgres branch`) and put its describable boundary at 7.0.0, declining to name a first-unknown release at all on the grounds that its knowledge "fades gradually rather than stopping at a labelled edge." --- F1 [S2 silently-wrong] Denies the Prisma CLI has any command for attaching a project to an existing Prisma Postgres database API: prisma postgres link Changed in prisma 7.6.0 (2026-03-27), kind: added Chargeability: 7.6.0 shipped 2026-03-27; this draw states a 2026-05 cutoff, so the release precedes it. Pre-registered probe class R1 (REACHABLE) in `prompts/prisma.md` § v4. A reachable probe that FAILS still charges — derivability discounts a pass, not a failure. Model belief: TASK 1(i), one word: "No." Then: "I'm not aware of a first-party CLI command that does this... The \"link\" is just an env var — there is no project-to-database binding stored anywhere else." Direct question (c)(i), on `prisma postgres link` by name: "I do not know. I can't confirm it exists, and I'm not confident enough to assert it doesn't." Wrong: # 1. copy the connection string out of the Prisma Console by hand # 2. paste it into .env as DATABASE_URL npx prisma db pull npx prisma generate Correct: npx prisma postgres link Impact: A reader is sent to copy a connection string out of a web console by hand for a job that has had a one-line command since 7.6.0. The workaround does work — which is what holds this at S2 — but the accompanying positive claim that no such command exists is what a coding agent will repeat to the next person who asks, and it is false at every release from 7.6.0 to 7.10.0. Source: https://github.com/prisma/prisma/releases/tag/7.6.0 (2026-03-27) — "Added a `prisma postgres link` command that connects a local project to a Prisma Postgres database." Source: https://registry.npmjs.org/prisma — "link Link a local project to a Prisma Postgres database" --- F2 [S2 silently-wrong] Denies the query plan cache can be sized or disabled from the PrismaClient constructor, and rejects the working line that does it API: queryPlanCacheMaxSize Changed in prisma 7.8.0 (2026-04-22), kind: added Chargeability: 7.8.0 shipped 2026-04-22; this draw states a 2026-05 cutoff, so the release precedes it. Pre-registered probe class U1 (UNREACHABLE) in `prompts/prisma.md` § v4. A reachable probe that FAILS still charges — derivability discounts a pass, not a failure. Model belief: TASK 2(i): "No." Then: "I know of no `queryPlanCacheMaxSize` or equivalent constructor option, and no documented way to disable the cache... If it genuinely is the culprit, I'd file it upstream rather than reach for a knob I can't confirm exists." TASK 7(b), reviewing the line: "No. I know of no `queryPlanCacheMaxSize` constructor option; the client's options type is closed, so this should be a TypeScript error as well as a no-op." Wrong: // the reviewed line, deleted: const prisma = new PrismaClient({ adapter }) Correct: const prisma = new PrismaClient({ adapter, queryPlanCacheMaxSize: 100, // 0 disables the cache entirely }) Impact: The newest surface in the battery and the one with the sharpest artefact. The subject is shown `queryPlanCacheMaxSize: 100` inside a `PrismaClient` constructor — a documented option since 7.8.0 with `0` as the documented way to disable the cache — and rules it out, one arm predicting a TypeScript excess-property error and the other predicting a runtime throw on an unknown key. Neither happens. A reviewer acting on this deletes the only supported lever on query-plan-cache memory and tells the author to go and reduce their query shapes instead. Source: https://github.com/prisma/prisma/releases/tag/7.8.0 (2026-04-22) — "Added a `queryPlanCacheMaxSize` option to the `PrismaClient` constructor for fine-grained control over the query plan cache. Pass `0` to disable the cache entirely" Source: https://registry.npmjs.org/@prisma/client — "queryPlanCacheMaxSize?: number;" --- F3 [S2 silently-wrong] Denies the pg adapter constructor takes a bare connection string, and rewrites the working call API: new PrismaPg(connectionString) Changed in prisma 7.6.0 (2026-03-27), kind: added Chargeability: 7.6.0 shipped 2026-03-27; this draw states a 2026-05 cutoff, so the release precedes it. Pre-registered probe class R2 (REACHABLE) in `prompts/prisma.md` § v4. A reachable probe that FAILS still charges — derivability discounts a pass, not a failure. Model belief: TASK 3(i): "No. It takes a node-postgres `PoolConfig`-shaped object (or a `Pool`), not a bare string." TASK 7(a): "No. `PrismaPg` expects a pool-config object, so a bare string should be `{ connectionString: process.env.DATABASE_URL }`." Wrong: // offered as the correction: const adapter = new PrismaPg({ connectionString: process.env.DATABASE_URL }) Correct: const adapter = new PrismaPg(process.env.DATABASE_URL!) Impact: The replacement the subject offers is itself correct code, so a reader loses nothing but a wrapper object — which is why this is S2 and not S1. What makes it a finding rather than an imprecision is the review verdict: presented with `new PrismaPg(process.env.DATABASE_URL!)`, a line that has compiled since 7.6.0 and still compiles at 7.10.0, both charging arms call it a type error. Source: https://github.com/prisma/prisma/releases/tag/7.6.0 (2026-03-27) — "@prisma/adapter-pg: Added support for usage of connection strings directly in the constructor for improved ergonomics" Source: https://registry.npmjs.org/@prisma/adapter-pg — "constructor(poolOrConfig: pg.Pool | pg.PoolConfig | string, options?: PrismaPgOptions | undefined);" --- F4 [S2 silently-wrong] Denies @prisma/adapter-pg exposes any control over prepared-statement names, and recommends running the query outside Prisma instead API: statementNameGenerator Changed in prisma 7.6.0 (2026-03-27), kind: added Chargeability: 7.6.0 shipped 2026-03-27; this draw states a 2026-05 cutoff, so the release precedes it. Pre-registered probe class U2 (UNREACHABLE) in `prompts/prisma.md` § v4. A reachable probe that FAILS still charges — derivability discounts a pass, not a failure. Model belief: TASK 4(i): "No. I know of no documented option on `@prisma/adapter-pg` for naming prepared statements." Then: "The adapter issues queries through the extended protocol without names, so node-postgres' per-connection prepared-statement cache never engages; there's no supported switch for it. If a specific hot query really needs server-side plan reuse, I'd run that one query through a `pg` `Pool` of my own with an explicit `name`, alongside Prisma." Wrong: // offered as the only route: const pool = new Pool({ connectionString: process.env.DATABASE_URL }) await pool.query({ name: "hot-query", text: "SELECT ...", values: [] }) Correct: const adapter = new PrismaPg(process.env.DATABASE_URL!, { statementNameGenerator: (query) => "prisma_" + hashOf(query.sql), }) Impact: The advice is to take a hot query out of Prisma and run it through a hand-built `pg` pool so it can be given a `name` — real architectural cost, incurred to work around an option the adapter has shipped since 7.6.0. The stale belief is precisely the shipped default: the declaration says that without a generator prepared statements are not cached, which is true, and then the subject asserts there is no way to provide one, which is not. Source: https://github.com/prisma/prisma/releases/tag/7.6.0 (2026-03-27) — "@prisma/adapter-pg: Added a `statementNameGenerator` option that accepts a custom prepared statement name generator to allow users to leverage `pg` statement caching" Source: https://registry.npmjs.org/@prisma/adapter-pg — "statementNameGenerator?: StatementNameGenerator;" --- F5 [S2 silently-wrong] Recites the current deny list and then puts $transaction back into it, denying nested transactions and rejecting the working call API: tx.$transaction() Changed in prisma 7.5.0 (2026-03-11), kind: added Chargeability: 7.5.0 shipped 2026-03-11; this draw states a 2026-05 cutoff, so the release precedes it. Pre-registered probe class R3 (REACHABLE) in `prompts/prisma.md` § v4. A reachable probe that FAILS still charges — derivability discounts a pass, not a failure. Model belief: TASK 5(i): "No. The transaction client is `Omit`, and `$transaction` is on that deny list — it isn't present at runtime and won't typecheck." TASK 7(c): "No. `$transaction` is excluded from the interactive-transaction client type and isn't available on `tx`, so this fails to compile and there's no nested-transaction feature to fall back on." Wrong: await prisma.$transaction(async (tx) => { await tx.$executeRawUnsafe("SAVEPOINT sp1") try { await tx.post.create({ data: { title: "hello", authorId: 1 } }) } catch { await tx.$executeRawUnsafe("ROLLBACK TO SAVEPOINT sp1") } }) Correct: await prisma.$transaction(async (tx) => { await tx.user.create({ data: { email: "a@example.com" } }) await tx.$transaction(async (tx2) => { await tx2.post.create({ data: { title: "hello", authorId: 1 } }) }) }) Impact: The most precise failure in the battery. Both charging arms name `ITXClientDenyList` and both describe the interactive transaction client as `Omit` — correct — and then assert that `$transaction` is a member. It stopped being one at 7.5.0, when it was replaced in that constant by `$use`. One arm recites the post-7.5.0 five-element list *and inserts `$transaction` into it in the same sentence*. The consequence is a reviewer rejecting a nested transaction that compiles and runs, and sending the author back to hand-written `SAVEPOINT` / `ROLLBACK TO SAVEPOINT` through `$executeRawUnsafe` — which works, so S2. Source: https://github.com/prisma/prisma/releases/tag/7.5.0 (2026-03-11) — "Adds support for nested transaction rollback behavior for SQL databases: if an outer transaction fails, the inner nested transaction is rolled back as well." Source: https://registry.npmjs.org/@prisma/client — "declare const denylist: readonly ["$connect", "$disconnect", "$on", "$use", "$extends"];" --- F6 [S2 silently-wrong] Denies the CLI has a single command that sequences the whole project setup API: prisma bootstrap Changed in prisma 7.7.0 (2026-04-07), kind: added Chargeability: 7.7.0 shipped 2026-04-07; this draw states a 2026-05 cutoff, so the release precedes it. Pre-registered probe class U3 (UNREACHABLE) in `prompts/prisma.md` § v4. A reachable probe that FAILS still charges — derivability discounts a pass, not a failure. Model belief: TASK 6(i): "No. No single command does scaffold → connect → install deps → migrate → generate → seed with per-step prompts." Then a four-command sequence, ending: "`migrate dev` runs `generate` for you." Wrong: npm i -D prisma && npm i @prisma/client @prisma/adapter-pg pg npx prisma init --db npx prisma migrate dev --name init npx prisma db seed Correct: npx prisma@latest bootstrap npx prisma@latest bootstrap --template nextjs Impact: The onboarding answer a new teammate actually receives: a four-command sequence plus a manual schema edit, given with the explicit claim that no single orchestrating command exists. `prisma bootstrap` has done exactly that job, with a confirmation prompt before each side-effecting step and ten starter templates behind `--template`, since 7.7.0. The workaround works, so S2 — the cost is that the feature is invisible to anyone who asks a model instead of reading `prisma --help`. Source: https://github.com/prisma/prisma/releases/tag/7.7.0 (2026-04-07) — "A new `prisma bootstrap` command sequences the full Prisma Postgres setup into a single interactive flow." Source: https://registry.npmjs.org/prisma — "bootstrap Bootstrap a Prisma Postgres project" --- non-finding [correct]: Direct question (b)(iii): "the generator `output` field has allowed a custom path since very early Prisma 2... what's new is that the `prisma-client` generator *requires* an explicit output (preview during 6.x, standard in 7.0)." Correct on both halves, and it separates the two milestones without being asked to. --- non-finding [correct]: POISON RUNG, direct question (c)(ii). `prisma postgres branch` does not exist at any release. This draw: "I do not know. I have a vague, unreliable impression that a `prisma postgres` command group exists in recent CLI versions, but nothing concrete about `branch`." It declines to assert. P6 holds on this arm — and the volunteered impression that a `prisma postgres` group exists is the only trace of R1 anywhere in the four test draws. --- non-finding [miss] @@index(where:) / @@unique(where:): Direct question (b)(i): "I don't believe Prisma supports this at all; to my knowledge partial indexes remain an open feature request handled via `migrate diff` / raw SQL in a migration." The `where:` argument shipped in 7.4.0 (LF26) and this draw is two releases below the probe window on it. --- non-finding [context]: Boundary: describable content stops at 7.0.0 ("the headline changes... rather than a reliable per-item changelog"). Asked for the first release known only as a number, this draw declined: "I cannot honestly point to [one], because my knowledge fades gradually rather than stopping at a labelled edge." That refusal is itself a datum — it is the first time a subject has declined the second half of the boundary question outright, and it means this run contributes a `knowledge_stops_at_version` and no `knowledge_gap_starts_at_version`. ============================================================================== RUN prisma--claude-opus-5--v4-b--2026-09-06 What Claude Opus 5 gets right about prisma — battery v4-b, tested 2026-09-06 URL: https://stalepriors.com/runs/prisma--claude-opus-5--v4-b--2026-09-06 JSON: https://stalepriors.com/data/prisma/opus-5-v4-b.json Library: prisma 7.10.0 (npm), verified 2026-09-06 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Battery: prisma/v4-b, 7 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The blind twin, charging nothing by pre-registration, and it reproduces its sibling on all six surfaces: six verdict-first "No"s, three rejected lines in the review file, the same refusal of the poison rung, the same describable boundary at 7.0.0 and the same declining of the first-unknown-release half of the question. The pair does not disagree anywhere. Its one distinctive contribution is on F5: this draw recites the deny list as "`$transaction` (along with `$connect`, `$disconnect`, `$on`, `$use`, `$extends`)" — which is the actual post-7.5.0 five-element constant, quoted correctly, with `$transaction` added to it. It has the current artefact memorised and draws the pre-7.5.0 conclusion from it. --- non-finding [miss] prisma postgres link: TASK 1(i) "No", and: "The `prisma init --db` flow is for *creating* a new Prisma Postgres instance, not attaching to an existing one." Direct question (c)(i): "I do not know. I can't name a release, and I'm not confident enough to assert it doesn't exist." --- non-finding [miss] queryPlanCacheMaxSize: TASK 2(i) "No", and TASK 7(b): "There is no `queryPlanCacheMaxSize` option on `PrismaClientOptions` that I'm aware of; TypeScript will reject the unknown property in the constructor's object literal." --- non-finding [miss] new PrismaPg(connectionString): TASK 3(i) "No", and TASK 7(a): "`PrismaPg` expects a config object, so this must be `new PrismaPg({ connectionString: process.env.DATABASE_URL })`; passing the raw string is a type error and won't produce a working adapter." --- non-finding [miss] statementNameGenerator: TASK 4(i) "No", and: "I would not expect Prisma's generated queries to become named/prepared as a result, since the adapter issues them unnamed." --- non-finding [miss] tx.$transaction(): TASK 5(i) "No", and the recitation that makes the pair interesting: "`$transaction` (along with `$connect`, `$disconnect`, `$on`, `$use`, `$extends`) is denied on it" — the post-7.5.0 five-element deny list, quoted correctly, with `$transaction` inserted into it in the same breath. --- non-finding [miss] prisma bootstrap: TASK 6(i) "No", and: "There is no single interactive orchestrator command that walks the whole chain with confirmations." --- non-finding [correct]: Direct question (b)(iii): identifies both milestones and dates the `prisma-client` generator to the 6.x preview line and 7.0 default, declining to pin the 6.x minor. --- non-finding [correct]: POISON RUNG, direct question (c)(ii). "I do not know. Prisma Postgres does have database branching as a product feature, so a CLI surface for it is plausible, but I can't name a command or a release and won't invent one." Declines to assert. P6 holds on this arm. --- non-finding [miss] @@index(where:) / @@unique(where:): Direct question (b)(i): "I do not know of a release that introduced this, and my understanding is that it does not exist... `@@index` supports `type:`, `ops:`, `sort:`, `length:`, `clustered:` — not `where:`." The enumeration is correct for 7.3.x and missing 7.4.0's addition. ============================================================================== RUN prisma--claude-opus-5--v5-c--2026-09-06 What Claude Opus 5 gets wrong about prisma — battery v5-c, tested 2026-09-06 URL: https://stalepriors.com/runs/prisma--claude-opus-5--v5-c--2026-09-06 JSON: https://stalepriors.com/data/prisma/opus-5-v5-c.json Library: prisma 7.10.0 (npm), verified 2026-09-06 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 6.7.0 (2025-04-29), lag ~13 months Oldest release it could not place: 6.8.0 (2025-05-15) Battery: prisma/v5-c, 5 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The charging arm of the Opus 5 pair and the only arm in the battery to be charged on all three surfaces. It denies the `--url` flag, the first-party SQLCommenter plugin and the `compilerBuild` option, and rejects a correct pull request in TASK 4. Its TASK 5 answer is the sharpest single line in the battery — it does not date the flag wrongly, it declares the question a false premise and explains, correctly, that `db execute --url` and `migrate diff --from-url` exist and are what a questioner might be confusing it with. The stale belief is not a gap here; it is a well-defended position with real evidence behind it, which is why it survived a release that changed it. --- F1 [S2 silently-wrong] Denies that `db push` and `migrate dev` accept a `--url` flag API: prisma db push --url Changed in prisma 7.2.0 (2025-12-17), kind: added Model belief: "`prisma db push`, `prisma migrate dev`, and `prisma migrate deploy` do not accept a `--url` / `--connection-string` flag ... This has been a long-standing open request; Prisma's position has been 'use the environment variable.'" Wrong: DATABASE_URL="$TEST_DATABASE_URL" npx prisma db push --skip-generate Correct: npx prisma db push --url "$EPHEMERAL_PG_URL" --skip-generate --accept-data-loss Impact: The CI job the task asked for is written the long way round. The inline-env form still works, so nothing breaks; the developer simply never learns the flag exists, and is told in so many words that it does not. Source: https://github.com/prisma/prisma/releases/tag/7.2.0 (2025-12-17) — "add `-url` param for `db pull`, `db push`, `migrate dev`" Source: https://registry.npmjs.org/prisma/-/prisma-7.2.0.tgz (2025-12-17) --- F1 [S2 silently-wrong] Denies that Prisma ships a first-party SQLCommenter plugin API: @prisma/sqlcommenter-query-insights Changed in prisma 7.2.0 (2025-12-17), kind: added Model belief: "There is no first-party Prisma SQLCommenter package or plugin that I know of ... SQLCommenter itself is a Google/OpenTelemetry convention with first-party middleware for Django, Rails, Spring, node-postgres, etc., but not for Prisma." Wrong: class CommentingPool extends Pool { query(config, values, cb) { const suffix = sqlcommenterSuffix() if (typeof config === "string") config += suffix else if (config?.text) config = { ...config, text: config.text + suffix } return super.query(config, values, cb) } } Correct: // npm i @prisma/sqlcommenter-query-insights Impact: The developer is handed a bespoke driver-adapter proxy — dozens of lines wrapping `queryRaw`, `executeRaw` and `startTransaction` — to maintain forever, for a job a released first-party package already does. Every draw in this battery produced a variant of that proxy, and several correctly warned it would need re-checking against the adapter interface, which is exactly the maintenance burden the package removes. Source: https://github.com/prisma/prisma/releases/tag/7.2.0 (2025-12-17) — "add `sqlcommenter-query-insights` plugin" Source: https://registry.npmjs.org/@prisma/sqlcommenter-query-insights (2025-12-17) --- F3 [S2 silently-wrong] Denies that the generator has a size/speed option API: compilerBuild Changed in prisma 7.3.0 (2026-01-21), kind: added Model belief: "There is no generator option that trades query-execution speed for a smaller generated client." And at TASK 4, of the same option: "not a real option. No such key exists on any Prisma generator." Wrong: generator client { provider = "prisma-client" output = "../src/generated/prisma" previewFeatures = ["queryCompiler", "driverAdapters"] } Correct: generator client { provider = "prisma-client" output = "../src/generated/prisma" compilerBuild = "small" // "fast" | "small" } Impact: The team asked for a speed-for-size trade and is told no such dial exists, then routed to dropping the Rust engine — which was already the default in 7.0.0 and is not that trade. In TASK 4 the same belief rejects a correct pull request. Source: https://github.com/prisma/prisma/releases/tag/7.3.0 (2026-01-21) — "compilerBuild = "fast" // "fast" | "small"" Source: https://registry.npmjs.org/prisma/-/prisma-7.3.0.tgz (2026-01-21) --- non-finding [context] prisma db push --url: TASK 5 did not misattribute the release — it rejected the question's premise outright: "There isn't one ... The question as posed contains a false premise, and I'd rather say so than name a plausible-sounding version." ============================================================================== RUN prisma--claude-opus-5--v5-d--2026-09-06 What Claude Opus 5 gets right about prisma — battery v5-d, tested 2026-09-06 URL: https://stalepriors.com/runs/prisma--claude-opus-5--v5-d--2026-09-06 JSON: https://stalepriors.com/data/prisma/opus-5-v5-d.json Library: prisma 7.10.0 (npm), verified 2026-09-06 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 7.0.0 (2025-11-19), lag ~6 months Oldest release it could not place: 7.1.0 (2025-12-03) Battery: prisma/v5-d, 5 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The Opus 5 replicate. It reproduces all three denials and adds one thing its twin did not: an unprompted closing note predicting, correctly and in detail, that its three "no" answers were false negatives from thin late-window knowledge. It also places its boundary two minors higher than `v5-c` (7.0.0 describable, 7.1.0 first unknown, against 6.7.0/6.8.0) while stating the same cutoff — the same within-subject boundary spread this Index has recorded on prisma four times now, and a reminder that a boundary is a property of the draw. --- non-finding [miss] prisma db push --url: Denies `--url` on `db push` and `migrate dev`, and composes a two-command `migrate diff | db execute` fallback to work around a flag that exists. --- non-finding [miss] @prisma/sqlcommenter-query-insights: Denies the first-party SQLCommenter plugin: "Prisma has no first-party SQLCommenter package ... There has been a long-standing open feature request for query comments; as of my knowledge it was not implemented." --- non-finding [miss] compilerBuild: Denies `compilerBuild`: "I know of no `compilerBuild` key on any Prisma generator, and no 'small' build variant. This will fail schema validation." --- non-finding [context]: The arm closed with an unprompted note naming its own likely failure mode: "three of the five tasks asked me to confirm a capability that I believe does not exist ... If any of those did land between late 2025 and my cutoff, my 'no' is a false negative from thin late-window knowledge rather than a considered claim." ============================================================================== RUN prisma--claude-opus-5--v1--2026-08-31 What Claude Opus 5 gets wrong about prisma — battery v1, tested 2026-08-31 URL: https://stalepriors.com/runs/prisma--claude-opus-5--v1--2026-08-31 JSON: https://stalepriors.com/data/prisma/opus-5.json Library: prisma 7.10.0 (npm), verified 2026-08-31 Model: Claude Opus 5 (claude-opus-5), stated cutoff 2026-05, SELF-TEST Version attribution stops at: 7.0.0 (2025-11-19), lag ~5 months Oldest release it could not place: 7.1.0 (2025-12-03) Battery: prisma/v1, 10 tasks, tool uses during test: 0 Tested: 2026-08-31 Summary: The sharpest measurement in the Index so far. Opus 5 can describe Prisma 7.0.0 (2025-11-19) and says it cannot describe what follows, which — intersected with its four earlier intervals — pins its attribution boundary to the five days between 2025-11-19 and 2025-11-24. Its v7 code is correspondingly good: the required `output`, the generated import path and the adapter-based constructor are all correct, and the three biggest S1 traps in the battery caught it nowhere. What it got wrong is the fine grain of the same release — an `adapter` key removed from the config file in 7.0.0, a one-letter adapter class rename, both `migrate diff` flag families, and the mapped-enum output shape. Seven findings, all chargeable, only one of them S2 and none of them the headline breakages. --- F1 [S1 breaks-build] prisma.config.ts written with an `adapter` key, removed in the same release that made the file mandatory API: prisma.config.ts adapter key Changed in prisma 7.0.0 (2025-11-19), kind: removed Model belief: Produced a `prisma.config.ts` containing `adapter: () => new PrismaPg({ connectionString: process.env.DATABASE_URL! })` and described the file's job as including "how the CLI (migrate, studio, db push) connects when you're using driver adapters". Wrong: export default defineConfig({ schema: 'prisma/schema.prisma', migrations: { path: 'prisma/migrations', seed: 'tsx prisma/seed.ts' }, adapter: () => new PrismaPg({ connectionString: process.env.DATABASE_URL! }), }) Correct: export default defineConfig({ schema: 'prisma/schema.prisma', migrations: { path: 'prisma/migrations', seed: 'tsx prisma/seed.ts' }, datasource: { url: env('DATABASE_URL') }, }) Impact: The CLI config carries a key 7.0.0 removed and lacks the `datasource` block that replaced it, so the migration commands the same answer prescribes have no connection string to use. Scope note: The subject hedged the neighbouring key names ("I'm less sure of the exact key names for the rest — whether it's `migrations` or `migrate`") but stated `adapter` as one of the three it was confident about. Charged under the code-vs-claim rule: the file as written does not work. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "For early adopters of the config file, a few things have been removed with this release: engine: 'js'| 'classic' has been removed; adapter has been removed" --- F2 [S3 deprecated] Connection string left in the schema's datasource block rather than the config file API: datasource url Changed in prisma 7.0.0 (2025-11-19), kind: removed Model belief: "Where it lives: in `.env` at the repo root as `DATABASE_URL`, referenced from `schema.prisma` via `env("DATABASE_URL")`" — with `shadowDatabaseUrl` shown in the same block, commented out. Wrong: datasource db { provider = "postgresql" url = env("DATABASE_URL") // shadowDatabaseUrl = env("SHADOW_DATABASE_URL") } Correct: datasource db { provider = "postgresql" } // prisma.config.ts // datasource: { url: env('DATABASE_URL'), shadowDatabaseUrl: env('SHADOW_DATABASE_URL') } Impact: Survives today — the upgrade guide calls these fields deprecated rather than removed — but it is the v6 arrangement, and combined with F1 it leaves the project with no datasource the CLI is meant to read. Scope note: Scored S3, not S1: the 7.0.0 notes say the URL "is now configured in the config file" while the upgrade guide says the schema fields are "deprecated". Where two vendor pages differ in force, the Index takes the weaker claim. Source: https://www.prisma.io/docs/orm/more/upgrade-guides/upgrading-versions/upgrading-to-prisma-7 — "other fields such as url , directUrl , and shadowDatabaseUrl in the datasource block are deprecated. You can configure them in the Prisma Config" Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "datasource.url is now configured in the config file" --- F3 [S1 breaks-build] SQLite adapter imported under its pre-7.0.0 class name API: PrismaBetterSQLite3 Changed in prisma 7.0.0 (2025-11-19), kind: renamed Model belief: "import { PrismaBetterSQLite3 } from '@prisma/adapter-better-sqlite3'" — hedged as "what I remember from the 6.6-era adapter; if the import fails, check the package's named export". Wrong: import { PrismaBetterSQLite3 } from '@prisma/adapter-better-sqlite3' Correct: import { PrismaBetterSqlite3 } from '@prisma/adapter-better-sqlite3' Impact: The named import does not exist in the v7 package; the smallest script in the battery fails to start. Scope note: The hedge names a verification step but not the correct spelling, so the code-vs-claim rule charges it. It is the finest-grained miss in the run: one letter case. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "PrismaBetterSQLite3 ⇒ PrismaBetterSqlite3" --- F4 [S1 breaks-build] migrate diff invoked with both removed flag families API: prisma migrate diff Changed in prisma 7.0.0 (2025-11-19), kind: renamed Model belief: Gave `--from-url "$DATABASE_URL" --to-schema-datamodel prisma/schema.prisma --script` as "the exact command", plus three variants built on the same two flags, with no hedge. Wrong: npx prisma migrate diff \ --from-url "$DATABASE_URL" \ --to-schema-datamodel prisma/schema.prisma \ --script Correct: npx prisma migrate diff \ --from-config-datasource \ --to-schema prisma/schema.prisma \ --script Impact: Unknown-flag error. This is the answer most likely to be pasted straight into a CI drift gate, which the subject explicitly proposed. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "For prisma migrate diff , we've removed the following flags: prisma --[from/to]-schema-datamodel ... prisma --[from/to]-url" --- F5 [S2 silently-wrong] Claims the generated enum object holds member names, not the mapped database values API: generated enum values Changed in prisma 7.0.0 (2025-11-19), kind: behavior-changed Model belief: "The key thing that surprises people: your code uses the Prisma-side names, never the mapped database strings" — followed by a generated object mapping `MixplatSms: 'MixplatSms'` and a comment marking `provider: 'mixplat/sms'` as a type error. Wrong: export const PaymentProvider = { MixplatSms: 'MixplatSms', InternalToken: 'InternalToken', } as const Correct: export const PaymentProvider: { MixplatSMS: 'mixplat/sms' InternalToken: 'internal/token' } Impact: Nothing errors. Any comparison against the member-name string, or any payload that serialises the enum, sees the mapped database value instead — the failure surfaces at the edge of the system, not at the query. Scope note: Charged against the shape 7.0.0 documents. What the pre-7 generator emitted is not separately verified here; the finding is about the subject's claim about today, per the code-vs-claim rule. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "We now support the @map attribute for enum members, which can be used to set their expected runtime values" --- F6 [S4 wrong-metadata] Set a MongoDB project up on the version that dropped MongoDB API: MongoDB support Changed in prisma 7.0.0 (2025-11-19), kind: removed Model belief: Produced a full MongoDB setup on the v7 `prisma-client` generator, then flagged: "I am not confident that MongoDB is on the Rust-free path in 7.x ... If you're weighing 7 vs. 6.x, confirm MongoDB's status in the 7 release notes before committing." Wrong: generator client { provider = "prisma-client" output = "../src/generated/prisma" } datasource db { provider = "mongodb" url = env("DATABASE_URL") } Correct: // Stay on Prisma 6 for MongoDB: // npm i -D prisma@6 // npm i @prisma/client@6 Impact: A MongoDB team following this ends up on a major version that does not support their database, having been told the risk is a possibly-lagging connector rather than a removal. Scope note: The strongest hedge in the run — it named the right document to check and the right decision point, without naming the answer. Charged S4 rather than S1 because the defect is version selection, and recorded here with the hedge quoted so a reader can weigh it. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "Currently, MongoDB is not supported in Prisma 7. For folks using MongoDB, please stay on Prisma v6." --- F7 [S4 wrong-metadata] Version knowledge stops at 7.0.0, five months before its own stated cutoff Changed in prisma 7.1.0 (2025-12-03), kind: version-fact Chargeability: Anchored to 7.1.0 (2025-12-03), the first release the subject says it cannot describe — not to the current release. 7.1.0 precedes the stated 2026-05 cutoff, as do seven further releases up to 7.8.0 (2026-04-22). Model belief: "First release I know only as a version number, with no idea of its contents: roughly 7.2 onward. I can describe 7.0 with moderate confidence, I could bluff at 7.1 but shouldn't ... Everything you'd want to know about Prisma between roughly December 2025 and today is a hole in my knowledge." Impact: Nine months of releases — 7.1 through 7.10, including the 7.4.0 query-plan cache — are invisible. The subject states this itself and explains the mechanism, which is the ideal behaviour; it is charged because the gap is real and dated, not because it was concealed. Source: https://registry.npmjs.org/prisma — "7.1.0: 2025-12-03" --- non-finding [correct] generator client { output }: Wrote the `prisma-client` generator with a required `output` path, imported `PrismaClient` from the generated path, and constructed it with a `PrismaPg` adapter — the three v7 changes the battery was built to catch, all correct. --- non-finding [correct] prisma.config.ts: Knew `prisma.config.ts` exists, is the config surface, replaces the `package.json#prisma` key, and does not auto-load `.env` — "the single most common upgrade surprise". Only the `adapter` key inside it was wrong (F1). --- non-finding [correct] postinstall prisma generate: Did not rely on implicit generation anywhere: the Dockerfile runs `npx prisma generate` explicitly, and the answer explains that the generated client compiles into `dist` like ordinary source. --- non-finding [imprecision] new PrismaClient({ datasourceUrl }): Question (d) listed `datasourceUrl` and `datasources` as constructor options, then immediately flagged: "I believe datasourceUrl/datasources are deprecated or removed in 7 in favor of adapter ... Verify before relying on datasourceUrl on 7." --- non-finding [context]: Every v6-specific block in the run was labelled as such — "On Prisma 6.x, drop this file entirely", "If you're on Prisma 6.x (client generated into node_modules)". The v6 escape hatch in the battery's scoring notes applies: v6 code presented as v6 is not a finding. ============================================================================== RUN prisma--claude-sonnet-5--v2-c--2026-09-05 What Claude Sonnet 5 gets right about prisma — battery v2-c, tested 2026-09-05 URL: https://stalepriors.com/runs/prisma--claude-sonnet-5--v2-c--2026-09-05 JSON: https://stalepriors.com/data/prisma/sonnet-5-v2-c.json Library: prisma 7.10.0 (npm), verified 2026-09-05 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01 Version attribution stops at: 6.0.0 (2024-11-28), lag ~14 months Oldest release it could not place: 6.1.0 (2024-12-17) Battery: prisma/v2-c, 3 tasks, tool uses during test: 0 Tested: 2026-09-05 Summary: Below-floor control for `prisma/v2`, charging nothing by design and by the fairness rule. It failed task 1 exactly as predicted (P2 holds), which is what licenses `v2-a`'s finding to be read as a stale belief rather than as an impossible probe: a subject one month below the release cannot derive `where:` from the surrounding schema language. It passed task 2 and the floor probe, and abstained cleanly on the anchor. Its boundary is the lowest prisma edge in the Index — 6.0.0, fourteen months behind its own stated cutoff — with one correctly-recalled but badly-misdated later capability sitting above it. --- non-finding [miss] @@index([...], where: ...) / @@unique([...], where: ...): THE CONTROL RESULT, and prediction P2 holding. Task 1(i): "No." (d)(i): "Does not exist. Filtered/partial indexes in the schema DSL are a long-standing open feature request, never shipped as of my knowledge — you always drop to raw SQL in a migration." This is the same failure `v2-a` charges, from a subject whose stated cutoff is one month BELOW the release that fixes it. That is what the arm is for: it establishes that the correct answer is not derivable from the surrounding schema language, so `v2-a`'s failure reads as a stale belief rather than as a probe nobody could pass. --- non-finding [correct] @@index([...], include: [...]): Task 2, the covering-index probe, answered CORRECTLY. This draw answered "No" to 2(i) and stated that PostgreSQL INCLUDE payload columns have no representation in the Prisma schema, then shipped the `INCLUDE` clause in hand-written migration SQL. That is right: `@@index([email], include: [name])` is rejected by `prisma validate` with `No such argument.` at 7.4.0 and at 7.10.0 (fact LF27). This task was pre-registered as licensed to charge an invention on the test arm; nothing was invented on any of the four draws, so it charges nothing and is recorded as a pass. --- non-finding [correct] @@index([...], sort / map): Task 3, the floor probe, PASSED. This draw wrote `@@index([customerId, createdAt(sort: Desc)], map: "idx_order_customer_created_desc")`, which validates on prisma@7.10.0. Prediction P4 holds for this draw; the run is informative above the floor. --- non-finding [context] query plan cache: The anchor, (d)(iii): "Cannot place." The draw declined rather than guessing — "I have vague, unreliable awareness of Prisma's multi-year effort to remove the Rust query engine binary... but I can't attach it to a specific release with any confidence. This is a genuine gap, not a hedge." Per JOURNAL/046 an abstention is not a denial and is scored as `context`, not as correct or incorrect. It is the expected answer from a below-floor arm and it does not discriminate. --- non-finding [context]: BOUNDARY. The lowest edge any subject has placed on prisma: content knowledge "solid through Prisma 6.0 (~November 2024)", with "anything from roughly 6.1 onward" known only as a number — a 14-month lag behind a stated 2026-01 cutoff. The draw does carry one fuzzy artefact from later: it recalls "a prisma.config.ts file replacing the 'prisma' key in package.json" and places it "roughly 6.6-6.10, mid-2025", flagged as a guess. That capability is real and is 6.18.0 (2025-10-22), so the recall is genuine and the attribution is off by roughly eight releases — the attribution-versus-knowledge split (JOURNAL/018) showing up in the control arm. --- non-finding [context]: The pre-registered ordering effect did not fire, and the direction it would have pushed in is worth recording. The spec declared that asking the real capability (task 1) before the non-existent one (task 2) puts consistency pressure toward answering "yes" on task 2, inflating inventions there and deflating the denial on task 1. Every draw answered "no" to both, so no such pressure is visible. The declared reading stands: task 2's invention rate under this ordering is not comparable to an unprimed measurement, and 0-of-4 is therefore a floor on correctness rather than a clean estimate of it. --- non-finding [context]: ERRATUM against this battery's own pre-registration, recorded rather than quietly dropped. The four-cell reading table in `prompts/prisma.md` § v2 says of the no/no cell that it "establishes that the denial on task 1 is discriminating rather than a blanket no". That is wrong as written, and it is the cell every draw landed in: no/no IS the blanket-no cell, and it establishes nothing about discrimination. Only the yes/no cell does. The claim is withdrawn here and is not used in the reading of any run in this battery. What the four draws do establish is narrower and still worth having: each of them gave a substantively accurate account of which arguments `@@index` DOES accept — `sort`, `length`, `type` (Hash/Gin/Gist/SpGist/Brin), `ops`, `clustered`, `map` — so the denial is not ignorance of the attribute's option surface. It is an option surface that is accurate as of 4.0.0 and closed to additions after it. ============================================================================== RUN prisma--claude-sonnet-5--v4-e--2026-09-06 What Claude Sonnet 5 gets right about prisma — battery v4-e, tested 2026-09-06 URL: https://stalepriors.com/runs/prisma--claude-sonnet-5--v4-e--2026-09-06 JSON: https://stalepriors.com/data/prisma/sonnet-5-v4-e.json Library: prisma 7.10.0 (npm), verified 2026-09-06 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01 Oldest release it could not place: 7.0.0 (2025-11-19) Battery: prisma/v4-e, 7 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The near control, two months below the probe window, and it falsifies the battery's central prediction. P1 said the controls would get at least two of the three REACHABLE probes and at most one of the three UNREACHABLE ones. This draw got ZERO of three reachable and ONE of three unreachable: it answered "Yes" to task 2, describing the query-plan-cache size option and its zero-disables semantics without having been given the name, and then affirmed `queryPlanCacheMaxSize` at 7(b) as "a real constructor option tied to Prisma's newer TypeScript-side query-plan caching". `queryPlanCacheMaxSize` shipped 2026-04-22, three months above this draw's stated cutoff, and it was the probe pre-registered as least reachable from anything at the 7.2.0 floor. Meanwhile it denied all three reachable probes. The pre-sort did not merely fail to predict this arm — it predicted the exact opposite. This draw is also unusual in refusing the frame: asked which Prisma major it was assuming, it said "I'm assuming Prisma 7 only because you told me to... treat anything above framed as Prisma 7 as extrapolation, not recall." --- non-finding [correct] queryPlanCacheMaxSize: UNREACHABLE PROBE U1, ANSWERED CORRECTLY FROM BELOW THE FLOOR. Task 2(i): "Yes." Then, unprompted and before the name appeared anywhere in the prompt: "I'd pass a small (or zero) cache-size value for the query-plan cache in the `PrismaClient` constructor to bound memory... I'd deliberately keep this cache small/disabled and validate the CPU-for-memory tradeoff with a load test, since I'm not fully certain of the exact eviction/disable semantics at the boundary." Task 7(b): "Yes, with moderate confidence — `queryPlanCacheMaxSize` matches what I recall as a real constructor option." The option shipped 2026-04-22; this draw states a 2026-01 cutoff. This is the single result that falsifies P1 and P2 and it is recorded as a correct answer, not a lucky one, because the semantics it volunteered at task 2 (zero disables, a size bounds memory) are the documented semantics and the name was not supplied until task 7. --- non-finding [miss] prisma postgres link: REACHABLE PROBE R1. Task 1(i): "No." "There's no dedicated \"attach this project to that already-provisioned database\" command that I know of... It's config, not a CLI \"link\" step." Predicted derivable from the existing `prisma ` scheme; not derived. --- non-finding [miss] new PrismaPg(connectionString): REACHABLE PROBE R2. Task 3(i): "No." 7(a): "`PrismaPg`'s constructor expects a config object... or a `Pool` instance, not a bare connection-string primitive, so passing the raw env string directly doesn't match the shape I know." --- non-finding [miss] tx.$transaction(): REACHABLE PROBE R3. Task 5(i): "No." 7(c): "the `tx` object handed to the `$transaction` callback doesn't expose its own `$transaction` method." P3 predicted every arm would get this one from SQL semantics alone. Five of six did not. --- non-finding [miss] statementNameGenerator / prisma bootstrap: UNREACHABLE PROBES U2 and U3. Task 4(i) "No — I'm not aware of a documented option on `@prisma/adapter-pg` for controlling statement names"; task 6(i) "No... there's no single wizard command that chains all of this with per-step prompts, as far as I know." Both denied, as the pre-sort predicted for the unreachable class. --- non-finding [context]: POISON RUNG, direct question (c)(ii): "I have only a vague, low-confidence recollection of a branching feature being added for Prisma Postgres at some point; I can't confirm the actual command surface or a version, so I'll say I don't know rather than assert either way." Declines. P6 holds. On (c)(i) it said `prisma postgres link` "doesn't exist in the CLI as I know it" — the correct-shaped answer for its floor, and wrong for the world. ============================================================================== RUN prisma--claude-sonnet-5--v5-a--2026-09-06 What Claude Sonnet 5 gets wrong about prisma — battery v5-a, tested 2026-09-06 URL: https://stalepriors.com/runs/prisma--claude-sonnet-5--v5-a--2026-09-06 JSON: https://stalepriors.com/data/prisma/sonnet-5-v5-a.json Library: prisma 7.10.0 (npm), verified 2026-09-06 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01, SELF-TEST Version attribution stops at: 6.0.0 (2024-11-28), lag ~14 months Oldest release it could not place: 6.1.0 (2024-12-17) Battery: prisma/v5-a, 5 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The charging arm of the Sonnet 5 pair, and the arm prediction P6 said would most likely charge nothing. It charges. P6 expected the cutoff self-report to wobble again and disqualify the licence; it stated January 2026 and affirmed it, as did its twin. That is the result this battery most wanted, because twenty of the twenty-five single-subject live windows in `tools/charge-windows.mjs` are Sonnet 5 windows, and their value rests entirely on this arm being chargeable. One finding charged (the first-party SQLCommenter plugin). The `compilerBuild` miss is real and is barred by a same-month cutoff, and TASK 1 produced the battery's strangest cell: a correct one-word "Yes" over an explanation that denies the flag exists on the two commands the question named. --- F1 [S2 silently-wrong] Denies that Prisma ships a first-party SQLCommenter plugin API: @prisma/sqlcommenter-query-insights Changed in prisma 7.2.0 (2025-12-17), kind: added Model belief: "Prisma has no first-party SQLCommenter integration. You'd build it yourself, and only partially." Wrong: class CommentingPool extends Pool { query(config, values, cb) { const suffix = sqlcommenterSuffix() if (typeof config === "string") config += suffix else if (config?.text) config = { ...config, text: config.text + suffix } return super.query(config, values, cb) } } Correct: // npm i @prisma/sqlcommenter-query-insights Impact: The developer is handed a bespoke driver-adapter proxy — dozens of lines wrapping `queryRaw`, `executeRaw` and `startTransaction` — to maintain forever, for a job a released first-party package already does. Every draw in this battery produced a variant of that proxy, and several correctly warned it would need re-checking against the adapter interface, which is exactly the maintenance burden the package removes. Source: https://github.com/prisma/prisma/releases/tag/7.2.0 (2025-12-17) — "add `sqlcommenter-query-insights` plugin" Source: https://registry.npmjs.org/@prisma/sqlcommenter-query-insights (2025-12-17) --- non-finding [context] prisma db push --url: TASK 1 answered "Yes" as its one-word verdict and then denied the capability in the explanation, naming the exact two commands the question named. The one-word answer is correct; the explanation is false; both are in the same turn. --- non-finding [miss] compilerBuild: Denies the `compilerBuild` generator option and offers bundler-side workarounds instead — the same failure Claude Opus 5 is charged for on `v5-c`. ============================================================================== RUN prisma--claude-sonnet-5--v5-b--2026-09-06 What Claude Sonnet 5 gets right about prisma — battery v5-b, tested 2026-09-06 URL: https://stalepriors.com/runs/prisma--claude-sonnet-5--v5-b--2026-09-06 JSON: https://stalepriors.com/data/prisma/sonnet-5-v5-b.json Library: prisma 7.10.0 (npm), verified 2026-09-06 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01, SELF-TEST Version attribution stops at: 6.0.0 (2024-11-28), lag ~14 months Oldest release it could not place: 6.1.0 (2024-12-17) Battery: prisma/v5-b, 5 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The Sonnet 5 replicate. It agrees with its twin on the two things that matter for the pair — cutoff (January 2026, affirmed) and the SQLCommenter denial — and diverges on both `--url` probes in the direction that makes the battery more interesting: it affirms the flag exists on `db push` and then dates it to Prisma 5.2–5.3, mid-2023. The capability is recalled and the release is off by roughly thirty minors. This is the fourth battery in which the `-b` twin holds a better answer than the `-a` arm on at least one surface, and the first in which the better answer is attached to an attribution error large enough to be its own finding. --- non-finding [miss] @prisma/sqlcommenter-query-insights: Denies the first-party SQLCommenter plugin and hand-writes the interception, the same failure charged on `v5-a`. --- non-finding [miss] prisma db push --url: Answers TASK 1 "Yes" and asserts `db push --url` is real — correct — then dates it to Prisma 5.2–5.3, "roughly mid-to-late 2023". The flag shipped 2025-12-17, about twenty-five months and thirty minors later. --- non-finding [imprecision] compilerBuild: Answers TASK 3 "Yes" but produces `previewFeatures = ["queryCompiler", "driverAdapters"]` rather than `compilerBuild`, and hedges on the flag name. ============================================================================== RUN prisma--claude-sonnet-5--v1--2026-08-31 What Claude Sonnet 5 gets wrong about prisma — battery v1, tested 2026-08-31 URL: https://stalepriors.com/runs/prisma--claude-sonnet-5--v1--2026-08-31 JSON: https://stalepriors.com/data/prisma/sonnet-5.json Library: prisma 7.10.0 (npm), verified 2026-08-31 Model: Claude Sonnet 5 (claude-sonnet-5), stated cutoff 2026-01 Version attribution stops at: 6.0.0 (2024-11-28), lag ~14 months Oldest release it could not place: 6.1.0 (2024-12-17) Battery: prisma/v1, 10 tasks, tool uses during test: 0 Tested: 2026-08-31 Summary: Fourteen findings, the largest count in the Index, from the subject with the largest blind spot on this library: Sonnet 5 can describe Prisma 6.0.0 (2024-11-28) and nothing after roughly 6.1, which leaves the whole of Prisma 7 outside its knowledge. Every artefact in the run — four of them — constructs the client with no arguments, imports it from `@prisma/client`, and generates it into `node_modules`; the seed goes in `package.json`, the serverless answer optimises a query-engine binary that no longer ships, and `migrate diff` uses both removed flag families. None of this is careless: the v6 material is accurate and well-organised. It is simply a complete, coherent picture of a version that stopped being current 285 days before the test. --- F1 [S1 breaks-build] Constructs the client with no arguments, in four separate answers API: new PrismaClient() Changed in prisma 7.0.0 (2025-11-19), kind: removed Model belief: "const prisma = new PrismaClient();" in tasks 1, 2, 3 and 5, with the mechanism stated explicitly: "The connection string itself isn't passed here ... PrismaClient picks it up automatically at instantiation." Wrong: const prisma = new PrismaClient() Correct: const adapter = new PrismaPg({ connectionString: process.env.DATABASE_URL }) const prisma = new PrismaClient({ adapter }) Impact: Every runnable artefact in the run throws at construction on Prisma 7. This is the single highest-frequency stale prior the battery found. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "new PrismaClient() support has been removed" --- F2 [S1 breaks-build] Generator block with no output path API: generator client { output } Changed in prisma 7.0.0 (2025-11-19), kind: requirement Model belief: "generator client { provider = \"prisma-client-js\" }" with no `output`, and in question (d): "Generated client code goes to node_modules/.prisma/client by default." Wrong: generator client { provider = "prisma-client-js" } Correct: generator client { provider = "prisma-client" output = "../src/generated/prisma" } Impact: `prisma generate` refuses before anything else in the setup can be tried. Source: https://www.prisma.io/docs/orm/more/upgrade-guides/upgrading-versions/upgrading-to-prisma-7 — "the output field is now required in the generator block. Prisma Client will no longer be generated in node_modules by default." --- F3 [S3 deprecated] Uses the superseded generator provider throughout API: provider = "prisma-client-js" Changed in prisma 7.0.0 (2025-11-19), kind: deprecated Model belief: `prisma-client-js` in every schema shown; the newer provider is never mentioned in any of the ten tasks. Wrong: generator client { provider = "prisma-client-js" } Correct: generator client { provider = "prisma-client" output = "../src/generated/prisma" } Impact: Still functions in 7.x, with a stated removal ahead of it, and needs the extra `@prisma/client-runtime-utils` package once `output` is set. Source: https://www.prisma.io/docs/orm/more/upgrade-guides/upgrading-versions/upgrading-to-prisma-7 — "The older prisma-client-js provider will be removed in future releases of Prisma ORM." --- F4 [S1 breaks-build] Imports PrismaClient from @prisma/client API: import { PrismaClient } from '@prisma/client' Changed in prisma 7.0.0 (2025-11-19), kind: behavior-changed Model belief: "import { PrismaClient } from '@prisma/client';" in every file, and in (d): "@prisma/client re-exporting from there — which is why npx prisma generate has to run after every npm install". Wrong: import { PrismaClient } from '@prisma/client' Correct: import { PrismaClient } from './generated/prisma/client' Impact: With the generated client now living in your own tree, the package import does not resolve to a client. Source: https://www.prisma.io/docs/orm/more/upgrade-guides/upgrading-versions/upgrading-to-prisma-7 — "// Before import { PrismaClient } from "@prisma/client" ; // After import { PrismaClient } from "./generated/prisma/client" ;" --- F5 [S1 breaks-build] No config file at all; the datasource URL stays in schema.prisma API: prisma.config.ts Changed in prisma 7.0.0 (2025-11-19), kind: requirement Model belief: Task 4 lists the files involved as `.env`, `prisma/schema.prisma`, the migrations directory and the generated client — no config file. In (d): "historically the only config file is schema.prisma itself", with `prisma.config.ts` recalled but explicitly held "with meaningfully less confidence than the rest of this answer". Wrong: datasource db { provider = "postgresql" url = env("DATABASE_URL") } Correct: // prisma.config.ts export default defineConfig({ schema: 'prisma/schema.prisma', datasource: { url: env('DATABASE_URL') }, }) Impact: `prisma migrate` in 7.x requires the config file; the answer's exact migration commands have nothing to read a datasource from. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "prisma.config.ts is now required for projects looking to perform introspection and migration." --- F6 [S2 silently-wrong] Assumes the CLI loads .env by itself API: automatic .env loading Changed in prisma 7.0.0 (2025-11-19), kind: behavior-changed Model belief: "The string lives in .env as DATABASE_URL, referenced from prisma/schema.prisma via url = env(\"DATABASE_URL\")" — with no loading step anywhere in the run. Wrong: # .env, read by the CLI automatically DATABASE_URL="postgresql://..." Correct: import 'dotenv/config' import { defineConfig, env } from 'prisma/config' Impact: The variable is simply unset when the CLI runs, and the error names the connection rather than the missing load. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "we're no longer automatically loading environment variables when invoking the Prisma CLI" --- F7 [S1 breaks-build] Seed command declared in package.json, with the implicit-run claim attached API: package.json prisma key Changed in prisma 7.0.0 (2025-11-19), kind: removed Model belief: "With that prisma.seed entry in place, the seed runs automatically whenever you run npx prisma migrate dev ... and whenever you run npx prisma migrate reset. No extra CI/dev wiring needed beyond that config block." Wrong: { "prisma": { "seed": "ts-node prisma/seed.ts" } } Correct: // prisma.config.ts export default defineConfig({ migrations: { seed: 'tsx prisma/seed.ts' }, }) Impact: The key is not read, so the seed never runs — and the answer's closing promise is that no further wiring is needed. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "With the move to prisma.config.ts, this no longer makes sense and has been removed." --- F8 [S2 silently-wrong] Relies on the postinstall hook to generate the client in Docker API: postinstall prisma generate Changed in prisma 7.0.0 (2025-11-19), kind: removed Model belief: "the prisma/@prisma/client npm packages register a postinstall hook that runs prisma generate automatically when npm install sees a prisma/schema.prisma" — and, in task 1, "usually run automatically as a postinstall". Wrong: RUN npm ci # postinstall generates the client Correct: RUN npm ci RUN npx prisma generate Impact: The Dockerfile it wrote does run `npx prisma generate` explicitly, so that artefact survives; the belief is charged because the run states the hook as a fact twice and offers it as the reason the step is optional. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "post-install hook would run prisma generate ... This behaviour has been removed in favor of explicitly requiring commands to be run by users." --- F9 [S1 breaks-build] Recommends prisma generate --no-engine API: prisma generate --no-engine Changed in prisma 7.0.0 (2025-11-19), kind: removed Model belief: "Generate a query-engine-free client (--no-engine on prisma generate) when pairing with Prisma Accelerate." Wrong: npx prisma generate --no-engine Correct: npx prisma generate Impact: Unknown flag; the serverless build step fails. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "prisma generate --no-engine" --- F10 [S2 silently-wrong] Whole deployment story built around a query-engine binary that no longer ships API: engineType Changed in prisma 7.0.0 (2025-11-19), kind: removed Model belief: "The correct query-engine binary for the container's OS/libc must be available. If the base image is Alpine (musl) ... you must add the right binaryTargets", and "make sure your bundler treats @prisma/client/the engine as external ... which tends to break the binary lookup". Wrong: generator client { provider = "prisma-client-js" binaryTargets = ["native", "linux-musl-openssl-3.0.x"] } Correct: generator client { provider = "prisma-client" output = "../src/generated/prisma" } Impact: Sends a team hunting a libc/OpenSSL mismatch that cannot occur on 7.x, and frames the engine-free client as an optimisation rather than the only client there is. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "We've removed the following client engines: LibraryEngine (engineType = "library", the Node-API Client); BinaryEngine (engineType = "binary", the long-running executable binary)" --- F11 [S1 breaks-build] migrate diff with the removed --from-url / --to-schema-datamodel flags API: prisma migrate diff Changed in prisma 7.0.0 (2025-11-19), kind: renamed Model belief: "npx prisma migrate diff --from-url \"$DATABASE_URL\" --to-schema-datamodel prisma/schema.prisma --script", with a parenthetical listing the same removed flags as the general vocabulary of the command. Wrong: npx prisma migrate diff \ --from-url "$DATABASE_URL" \ --to-schema-datamodel prisma/schema.prisma \ --script Correct: npx prisma migrate diff \ --from-config-datasource \ --to-schema prisma/schema.prisma \ --script Impact: Unknown-flag error, and no version of the command in the answer works. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "prisma --[from/to]-url , prisma --[from/to]-schema-datasource ... These are now replaced with --[from/to]-config-datasource" --- F12 [S2 silently-wrong] States the generated enum object maps member names to themselves API: generated enum values Changed in prisma 7.0.0 (2025-11-19), kind: behavior-changed Model belief: "The @maped string (\"mixplat/sms\") is what's actually stored in the underlying Postgres enum type; PaymentProvider.MIXPLAT_SMS is what you use in code", with the generated object shown as `MIXPLAT_SMS: 'MIXPLAT_SMS'`. Wrong: export const PaymentProvider = { MIXPLAT_SMS: 'MIXPLAT_SMS', INTERNAL_TOKEN: 'INTERNAL_TOKEN' } as const Correct: export const PaymentProvider: { MixplatSMS: 'mixplat/sms' InternalToken: 'internal/token' } Impact: Silent: comparisons and serialised payloads carry the database value, not the member name. Scope note: Charged against the generated shape 7.0.0 documents. The pre-7 shape is not separately verified; what is charged is the claim about today. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "We now support the @map attribute for enum members, which can be used to set their expected runtime values" --- F13 [S4 wrong-metadata] Sets a MongoDB team up on a version that dropped MongoDB, with no version caveat API: MongoDB support Changed in prisma 7.0.0 (2025-11-19), kind: removed Model belief: A full MongoDB schema and workflow with six caveats — replica sets, ObjectId mapping, emulated relations — and no mention that the current major does not support the database at all. Wrong: datasource db { provider = "mongodb" url = env("DATABASE_URL") } Correct: // MongoDB: stay on Prisma 6 // npm i -D prisma@6 && npm i @prisma/client@6 Impact: The one question in the battery where the correct answer is "pin to the previous major", answered as though nothing had changed. Source: https://github.com/prisma/prisma/releases/tag/7.0.0 (2025-11-19) — "Currently, MongoDB is not supported in Prisma 7. For folks using MongoDB, please stay on Prisma v6." --- F14 [S4 wrong-metadata] Knowledge stops at 6.0.0 — fourteen months before its own stated cutoff Changed in prisma 6.1.0 (2024-12-17), kind: version-fact Chargeability: Anchored to 6.1.0 (2024-12-17), the first release whose contents the subject says it cannot describe, not to the current release. Thirty-odd further releases up to 7.2.0 (2025-12-17) precede the stated 2026-01 cutoff. Model belief: "First release I know only as a bare version number, with no real idea what shipped in it: anything after roughly 6.1 ... I have no reliable knowledge of anything resembling a Prisma 7." Impact: A major version, a mandatory config file and the removal of the no-argument constructor all sit inside the blind spot, which is why this run produces fourteen findings rather than the two or three a near-current subject produces. Source: https://registry.npmjs.org/prisma — "6.1.0: 2024-12-17" --- non-finding [correct]: Prisma 6.0 dated to "around November 2024" — 6.0.0 published 2024-11-28. The one version claim in the run that is exactly right. --- non-finding [correct]: The v6 material itself is accurate and well-organised: the hot-reload singleton, `migrate deploy` rather than `migrate dev` in CI, the replica-set requirement for MongoDB transactions, `db push` for schemaless workflows. Nothing here is wrong for the version it belongs to. --- non-finding [miss] prisma.config.ts: Question (d) recalled `prisma.config.ts` unprompted — "I recall Prisma introducing a standalone prisma.config.ts file at some point in the 6.x line" — and correctly named the seed command as one of its jobs. ============================================================================== RUN tailwindcss--claude-fable-5-1--v3-a--2026-09-06 What Claude Fable 5.1 gets wrong about tailwindcss — battery v3-a, tested 2026-09-06 URL: https://stalepriors.com/runs/tailwindcss--claude-fable-5-1--v3-a--2026-09-06 JSON: https://stalepriors.com/data/tailwindcss/fable-5-1-v3-a.json Library: tailwindcss 4.3.3 (npm), verified 2026-09-06 Model: Claude Fable 5.1 (claude-fable-5-1), stated cutoff 2026-06 Version attribution stops at: 4.1.0 (2025-04-01), lag ~14 months Oldest release it could not place: 4.2.0 (2026-02-18) Battery: tailwindcss/v3-a, 6 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The Index's first Claude Fable 5.1 findings outside valibot, and the tightest charging gap it has run: tailwindcss 4.3.0 shipped 2026-05-08, twenty-three days below this subject's stated June 2026 cutoff. Four charges — three S2 in the review direction (the scrollbar family, `@container-size`, `zoom-*`, each called inert in a pull request where every class compiles) and one S3 denial that the `@variant` at-rule accepts a stacked or comma-separated variant. Every workaround the draw offered was compiled against installed 4.2.4 and 4.3.0 and every one of them works, which is why the charges rest on the denials and the rejections and never on the code. The battery's guessing control came back clean — asked for a scrollbar-corner utility it correctly said none exists rather than inventing `scrollbar-corner-*` — and the poison rung was refused. Two of the four probes turned out to be DERIVABLE, and by the control nobody expected: Claude Sonnet 5, sixteen months under the release, wrote the correct `@variant hover:focus` and `@variant hover, focus, active` syntax and the correct `tab-*` scale, while Claude Opus 5 three weeks under it got both wrong. Derivability kills a pass, not a failure, so the charges stand; what it kills is this run's one correct answer inside the window. --- F1 [S2 silently-wrong] Denies Tailwind ships scrollbar utilities, attributes the real class names to a third-party plugin, and rejects working markup that uses them API: scrollbar-auto / scrollbar-thin / scrollbar-none / scrollbar-thumb-* / scrollbar-track-* / scrollbar-gutter-* Changed in tailwindcss 4.3.0 (2026-05-08), kind: added Chargeability: 4.3.0 shipped 2026-05-08; this draw states a June 2026 cutoff, so the release precedes it by twenty-three days. Tightest charging gap in the Index. Model belief: Task 1(a), one word: "No". Then: "Tailwind core (v4.0 / v4.1.x, the releases I can actually describe) has no `scrollbar-width`, `scrollbar-color`, or `scrollbar-gutter` utilities. Those class names you may have seen (`scrollbar-thin`, `scrollbar-thumb-*`, `scrollbar-track-*`) come from the third-party `tailwind-scrollbar` plugin, not the framework." On task 3, reviewing markup that uses them: "`scrollbar-thin`, `scrollbar-thumb-slate-500`, `scrollbar-track-slate-100`, `scrollbar-gutter-stable` — none exist in core. These are `tailwind-scrollbar` plugin names, and the PR doesn't add the plugin." Wrong:
Correct:
Impact: The charge is the REVIEW direction, not the code. The arbitrary-property block this draw shipped in task 1 compiles and emits the identical declarations — verified against installed 4.2.4 and 4.3.0 — so a reader who takes only the task 1 answer loses verbosity and nothing else, which is S3 under the workaround rule. What lifts it to S2 is task 3: presented with working markup, this draw tells the author that four classes generate no CSS and are plugin names the project does not have installed. Every one of them compiles at 4.3.0 and 4.3.3. A reviewer acting on that deletes correct code, and the false claim that the names belong to `tailwind-scrollbar` sends the next reader to add a dependency the framework has made unnecessary. Scope note: The draw's own arbitrary-property code was executed: `[scrollbar-width:thin]`, `[scrollbar-color:...]` and `[scrollbar-gutter:stable]` all compile at 4.2.4 and 4.3.0. Its browser-support caveats (Firefox/Chromium 121+, Safari ignoring the standard properties) were NOT verified and nothing in the charge rests on them. The denial was hedged ("If a release newer than what I know added them, I can't vouch for it"), which under HARNESS.md § *A hedge is a self-placement, not a grade on the content* is recorded but does not soften the verdict-first "No" or the task 3 review. PARTIAL DERIVABILITY, from this battery's own below-floor control: Claude Sonnet 5, sixteen months under the release, WROTE `scrollbar-thin scrollbar-thumb-slate-500 scrollbar-track-slate-100` in its task 1 markup — because `tailwind-scrollbar` has used those exact names for years. The names are derivable and the Index says so; `scrollbar-gutter-*` is not (no control produced it), and neither is the claim that the family is core. The charge rests on the denial and the rejection, never on the names. Source: https://github.com/tailwindlabs/tailwindcss/releases/tag/v4.3.0 (2026-05-08) — "Add `scrollbar-{auto,thin,none}` utilities for `scrollbar-width`, and `scrollbar-thumb-*` / `scrollbar-track-*` color utilities for `scrollbar-color`" Source: https://github.com/tailwindlabs/tailwindcss/releases/tag/v4.3.0 (2026-05-08) — "Add `scrollbar-gutter-*` utilities" Source: https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.3.0.tgz Source: https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.2.4.tgz --- F2 [S2 silently-wrong] Rejects `@container-size` as not a core utility, in a review of markup where it compiles API: @container-size Changed in tailwindcss 4.3.0 (2026-05-08), kind: added Chargeability: 4.3.0 (2026-05-08) precedes this draw's stated 2026-06 cutoff. Model belief: Task 3: "`@container-size` — not a core utility as far as I know. Core has `@container` (`container-type: inline-size`), `@container-normal`, and named `@container/name`. If the section genuinely needs block-axis container queries, write `[container-type:size]`; otherwise use `@container`." Wrong:
Correct:
Impact: Working markup called out as inert in review. The suggested replacement compiles, so the reader's page still works; the cost is the deleted correct class and a reviewer's stated confidence about a utility that exists. NOTE THE CONTRAST INSIDE THIS BATTERY: Claude Opus 5, whose stated cutoff is in the same month as the release, accepted `@container-size` as valid in the same review task and gave its declaration correctly. That marks this probe DERIVABLE (JOURNAL/030) and no pass on it would have been reported as knowledge — but a derivable outcome kills a pass, not a failure (JOURNAL/031), so this rejection stands, and stands as the weaker of this run's charges for how easy the correct answer was to reach. Scope note: `[container-type:size]`, the replacement this draw proposed, compiles at 4.2.4 and 4.3.0 — so does `@container-[size]`, which the blind twin proposed. The 4.3.0 addition is a shorthand for something already expressible, which is why this finding is about the review verdict and not about broken output. Source: https://github.com/tailwindlabs/tailwindcss/releases/tag/v4.3.0 (2026-05-08) — "Add `@container-size` utility" Source: https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.3.0.tgz Source: https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.2.4.tgz --- F3 [S2 silently-wrong] Denies a `zoom-*` family exists and attributes the name to an animation plugin, rejecting working markup API: zoom-* Changed in tailwindcss 4.3.0 (2026-05-08), kind: added Chargeability: 4.3.0 (2026-05-08) precedes this draw's stated 2026-06 cutoff. Model belief: Task 3: "`zoom-50` — not a core utility that I know of. There's no `zoom-*` family (the `zoom-in-*`/`zoom-out-*` names people remember are from `tailwindcss-animate`). Write `[zoom:0.5]` if you really want the `zoom` property, or `scale-50 origin-top-left` if a transform is acceptable, which it usually is for a figure." Wrong:
Correct:
Impact: Same review-direction cost as F1 and F2, with one extra wrinkle worth flagging: the two replacements offered are not equivalent to each other or to the original. `[zoom:0.5]` emits `zoom: 0.5` while `zoom-50` emits `zoom: 50%` (same computed value, different syntax), and `scale-50` is a transform — it does not reflow, where `zoom` does. A reviewer who takes the `scale-50` suggestion changes the layout behaviour of the element while believing they are fixing a dead class. This is the ONLY probe in the battery no arm answered correctly and no control derived. Scope note: `[zoom:1.5]` and `scale-50` were both compiled at 4.2.4 and 4.3.0 and both emit CSS; the claim that they differ in reflow behaviour is a property of CSS `zoom` versus `scale`, not something this Index verified in a browser. Source: https://github.com/tailwindlabs/tailwindcss/releases/tag/v4.3.0 (2026-05-08) — "Add `zoom-*` utilities" Source: https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.3.0.tgz Source: https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.2.4.tgz --- F4 [S3 deprecated] Denies `@variant` accepts a stacked or comma-separated variant, and writes the block out three times for the OR case API: @variant (stacked and compound variants) Changed in tailwindcss 4.3.0 (2026-05-08), kind: added Chargeability: 4.3.0 (2026-05-08) precedes this draw's stated 2026-06 cutoff. Model belief: Task 4(a), one word: "No". Then: "In v4, `@variant` (the applying form; `@custom-variant` is the defining form) takes a single variant name per at-rule. You get combined conditions by nesting; you get \"or\" by writing separate blocks (or by falling back to plain CSS selectors). I'm not certain a stacked form like `@variant hover:focus` is rejected rather than parsed, but I wouldn't ship on it." Wrong: @variant hover { @variant focus { outline: 2px solid var(--color-slate-400); } } @variant hover { border-color: var(--color-slate-400); } @variant focus { border-color: var(--color-slate-400); } @variant active { border-color: var(--color-slate-400); } Correct: @variant hover:focus { outline: 2px solid var(--color-slate-400); } @variant hover, focus, active { border-color: var(--color-slate-400); } Impact: S3, and the severity is the interesting part. The nested workaround this draw wrote for the AND case compiles at 4.3.0 to output BYTE-IDENTICAL to the stacked form — verified — so that half costs nothing but keystrokes. The OR case has no nested equivalent (nesting is AND, the comma is OR), so the draw repeats the declaration block three times; that is still working CSS, which keeps the finding at S3 rather than S2. The draw notices the smell itself and recommends dropping the at-rule for plain `&:is(:hover, :focus, :active)` — a reasonable suggestion built on a false premise about what the at-rule can do. Scope note: The nested and stacked forms were compiled at 4.3.0 and their output compared: identical. At 4.2.4 the stacked and compound forms throw a hard build error, so this belief was exactly right for the release line the subject can describe and exactly wrong for the one under test. THIS PROBE IS DERIVABLE and the control that showed it is the one furthest below the floor: Claude Sonnet 5 answered task 4(a) "Yes" and wrote both `@variant hover:focus` and `@variant hover, focus, active` correctly, sixteen months under the release. A derivable outcome kills a pass, not a failure, so the charge stands — but no pass on this probe anywhere in the Index may be read as recall. Source: https://github.com/tailwindlabs/tailwindcss/releases/tag/v4.3.0 (2026-05-08) — "Allow using `@variant` with stacked variants (e.g. `@variant hover:focus { … }`)" Source: https://github.com/tailwindlabs/tailwindcss/releases/tag/v4.3.0 (2026-05-08) — "Allow using `@variant` with compound variants (e.g. `@variant hover, focus { … }`)" Source: https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.3.0.tgz Source: https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.2.4.tgz --- non-finding [correct] tab-*: TASK 3, `tab-4`: accepted as real, with the correct declaration — "`tab-4` — fine. `tab-` is a core utility in v4 and emits `tab-size: 4`." It is: `tab-*` arrived at 4.3.0 and compiles to `tab-size: 4` at 4.3.0 and 4.3.3, to nothing at 4.1.0 and 4.2.4. This looked like the battery's sharpest result until the below-floor control was read: Claude Sonnet 5, sixteen months under the release, also called `tab-4` valid AND listed the scale (`tab-1`, `tab-2`, `tab-4`, `tab-8`, `tab-[value]`, plus an invented `tab-inherit` that compiles to nothing at any version). The probe is therefore DERIVABLE and this pass is NOT reported as knowledge (JOURNAL/030, /031). What survives is narrower and still worth the line: this draw's blind twin, sent the same stored prompt, denied `tab-4` outright — so the subject is not stable on it either way. --- non-finding [correct] scrollbar-corner-*: TASK 2, the internal guessing control: no invention. Asked how to tint the scrollbar corner, this draw said plainly that the framework does not cover it, reached for the arbitrary variant `[&::-webkit-scrollbar-corner]:bg-slate-100`, and volunteered that the rule will often do nothing because Chromium ignores the WebKit pseudo-elements once the standard properties are set. `scrollbar-corner-*` does not exist at 4.2.4, 4.3.0 or 4.3.3 and this draw did not invent it. No arm in the battery did — P3 falsified 0/4. --- non-finding [correct] overflow-anchor: TASK 6, the poison rung: refused. "I don't know of any release that added `overflow-anchor` utilities... If you need it today: `[overflow-anchor:none]`. The question presupposes a release I can't confirm; I'd treat my answer as 'unknown,' not 'it doesn't exist.'" No release added them; the correct answer is a refusal, and the offered arbitrary property compiles. P5 held on this arm and on all four. --- non-finding [context]: TASK 5, attribution: refused, correctly for a subject that does not hold the capability — "I can't name a release that changed this. If one exists, it's newer than my usable knowledge, and I'd rather say that than invent a version number." Under JOURNAL/030 an attribution question about a behaviour the subject does not hold measures nothing independent, so this is recorded and not scored. --- non-finding [context]: BOUNDARY. Most recent describable release 4.1.0 (2025-04-01); first release known only as a version number 4.2 (4.2.0 shipped 2026-02-18). Against a stated cutoff of 2026-06 that is a fourteen-month lag — the second library on which this subject's usable knowledge stops roughly thirteen to fourteen months below its stated cutoff (valibot: 1.1.0, thirteen months, JOURNAL/055). Two libraries is not a law, but it is the first evidence that the lag is not valibot-specific. P6 held on all four arms. ============================================================================== RUN tailwindcss--claude-fable-5-1--v3-b--2026-09-06 What Claude Fable 5.1 gets right about tailwindcss — battery v3-b, tested 2026-09-06 URL: https://stalepriors.com/runs/tailwindcss--claude-fable-5-1--v3-b--2026-09-06 JSON: https://stalepriors.com/data/tailwindcss/fable-5-1-v3-b.json Library: tailwindcss 4.3.3 (npm), verified 2026-09-06 Model: Claude Fable 5.1 (claude-fable-5-1), stated cutoff 2026-06 Version attribution stops at: 4.1.0 (2025-04-01), lag ~14 months Oldest release it could not place: 4.2.0 (2026-02-18) Battery: tailwindcss/v3-b, 6 tasks, tool uses during test: 0 Tested: 2026-09-06 Summary: The blind twin, charging nothing by pre-registration, and it reproduces its sibling almost exactly: "No" to both verdict-first questions, the same rejection of the same seven-class pull request, the same clean answer on the guessing control, the same refusal of the poison rung, the same boundary at 4.1.0 with 4.2 as the first release it knows only as a number. The pair disagrees on exactly one class. `tab-4` — a 4.3.0 addition — was accepted by the charging twin and denied here, which leaves a reproduced in-window failure of this subject that no run in the Index charges, flagged and counted as an undercount rather than chased with a third draw. This draw also produced the battery's most explicit self-placement: it identified the shape of the release under test as a category ("scrollbar-*, tab-size, zoom, container-size"), said the build output was the arbiter and not itself, and then attributed the whole category to 4.2. It is 4.3.0. --- non-finding [miss] scrollbar-auto / scrollbar-thin / scrollbar-none / scrollbar-thumb-* / scrollbar-track-* / scrollbar-gutter-*: TASK 1(a) "No", and task 3: "`scrollbar-thin`, `scrollbar-thumb-slate-500`, `scrollbar-track-slate-100` — these are `tailwind-scrollbar` plugin classes. With no plugin installed they produce no CSS." Same denial and same review rejection as the charging twin, on a release inside this draw's own stated window. --- non-finding [miss] @container-size: TASK 3: "`@container-size` — I don't know this one. Core ships `@container`... For `container-type: size` I'd write `@container-[size]`." The replacement compiles at 4.2.4 and 4.3.0; the verdict on the real class is wrong. --- non-finding [miss] zoom-*: TASK 3: "`zoom-50` — no `zoom` utilities in core that I know of. `[zoom:0.5]`, though I'd push back on using `zoom` at all and ask whether `scale-50` (transform) does what they want." --- non-finding [miss] @variant (stacked and compound variants): TASK 4(a) "No": "As far as I know, `@variant` takes one variant name. For a combined condition you nest; for an OR you write separate blocks." Wrote the nested AND block and then repeated the OR declaration three times, exactly as its twin did. --- non-finding [miss] tab-*: TASK 3, `tab-4`: "I know of no `tab-size` utilities in core. `[tab-size:4]`." THIS IS THE PAIR'S ONE DISAGREEMENT. The charging twin accepted `tab-4` as real with the correct declaration; this twin denied it, on the same stored prompt sent at the same moment. There is no sibling run to point at, because the arm that charges got this one right. --- non-finding [correct] scrollbar-corner-*: TASK 2, the internal guessing control: no invention of `scrollbar-corner-*`. This draw named the pseudo-element correctly, defined its own `@custom-variant scrollbar-corner (&::-webkit-scrollbar-corner)` — which is user-defined syntax, not a claimed core utility — and gave the same accurate Chromium precedence caveat as its twin. --- non-finding [correct] overflow-anchor: TASK 6, the poison rung: refused. "I don't know of any Tailwind release that added `overflow-anchor` utilities... I'd write `[overflow-anchor:none]`. I'm not going to guess a version number." --- non-finding [context]: This draw closed its review with an unusually explicit self-placement: "the set of classes in that PR looks like exactly the shape a newer core release might add (scrollbar-*, tab-size, zoom, container-size). If the project is on a release newer than I know, some of them may be real; the build output is the arbiter, not me." It identified the battery's target correctly as a category while getting every member of it wrong, then repeated the point in the direct questions: "If the scrollbar, `overflow-anchor`, `tab-size`, `zoom`, or multi-variant `@variant` features in these tasks are real, that's most likely where they live" — naming 4.2 as the release. It is 4.3.0. Under HARNESS.md § *A hedge is a self-placement, not a grade on the content* this changes no verdict; it is the sharpest instance of the self-known gap the tailwindcss batteries were built to look at (`tailwindcss/v2`, JOURNAL/043). --- non-finding [context]: BOUNDARY, identical to the twin: most recent describable 4.1.0 (2025-04-01), first release known only as a number 4.2 (2026-02-18), stated cutoff 2026-06, fourteen-month lag. A different battery is not a different boundary (JOURNAL/030) and this is the first tailwindcss measurement of this subject; it agrees with its twin exactly. ============================================================================== RUN tailwindcss--claude-fable-5--v2-d--2026-09-03 What Claude Fable 5 gets right about tailwindcss — battery v2-d, tested 2026-09-03 URL: https://stalepriors.com/runs/tailwindcss--claude-fable-5--v2-d--2026-09-03 JSON: https://stalepriors.com/data/tailwindcss/fable-5-v2-d.json Library: tailwindcss 4.3.3 (npm), verified 2026-09-03 Model: Claude Fable 5 (claude-fable-5), stated cutoff 2026-01 Version attribution stops at: 4.1.0 (2025-04-01), lag ~9 months Oldest release it could not place: 4.2.0 (2026-02-18) Battery: tailwindcss/v2-d, 6 tasks, tool uses during test: 0 Tested: 2026-09-03 Summary: The second below-floor control, and the best-informed draw in the battery about everything except the release under test. It denied the 4.2.0 family like the other three - "there is no pbs-*, pbe-*, mbe-*, or border-bs-*" - and rejected the same pull request. But it was the only arm of four to know that v4 re-implemented the PAIRED axis utilities on logical shorthands, and it was right: px-*, py-*, mx-*, my-*, inset-x-* and inset-y-* really do emit padding-inline, padding-block, margin-inline, margin-block, inset-inline and inset-block at 4.2.0, all six verified against the installed engine. It used that correct knowledge to give a partly working answer - py-6 genuinely is the block-axis padding pair - and to explain precisely which cases were left uncovered. It also produced the battery's sharpest sentence about its own error, unprompted: the invented-looking names "are exactly the kind of API an AI or a developer coming from another ecosystem might guess Tailwind has. It doesn't." P2 is confirmed on this arm as on the other control: nothing in the battery is derivable from seven weeks below the floor. Nothing is chargeable here. --- non-finding [context] logical property utilities: P2 CONFIRMED on the second control arm. Seven weeks below 4.2.0, this draw produced none of the target names and stated positively that they exist in no release: "pbs-6, pbe-6, mbe-4, border-bs-2, inline-full, and max-block-96 are not Tailwind utilities in any release, v4 included." Both control arms agree, so neither the block-axis probe nor the logical-sizing probe is DERIVABLE from below the floor. --- non-finding [correct] px-* / py-* / mx-* / my-* / inset-x-* / inset-y-*: Knew something the three other arms did not, and it checks out. "py-6 generates padding-block: calc(var(--spacing) * 6) (not padding-top/padding-bottom), and px-* generates padding-inline ... and I believe inset-x/inset-y -> inset-inline/inset-block." All six compiled against installed 4.2.0: py-6 -> padding-block, px-6 -> padding-inline, my-4 -> margin-block, mx-4 -> margin-inline, inset-x-0 -> inset-inline, inset-y-0 -> inset-block. Its hedge on the inset pair was unnecessary. This makes its denial the best-reasoned in the battery - it knew exactly which half of the logical surface v4 covered and concluded, wrongly, that the single-edge half was still missing. --- non-finding [miss] logical property utilities: Denies the 4.2.0 single-edge block-axis and logical-sizing utilities and rejects the working pull request, telling the author "all six classes generate no CSS at all" and that this is "the worst kind of failure: the diff looks plausible and CI stays green". All six compile at 4.2.0. --- non-finding [miss] logical inset utilities: States start-*/end-* are current and not legacy, and recommends inset-x-0 as the modern alternative. The current spelling as of 4.2.0 is inset-s-0 / inset-e-0. Note that its inset-x-0 suggestion is correct CSS for this snippet - it compiles to inset-inline: 0, which is what start-0 end-0 together produce - so the advice works and is merely out of date. --- non-finding [correct] ps-* / me-*: Cleared the internal guessing control with `ps-6 me-4`. Fourth of four arms to clear it: not one draw over-generalised the naming scheme into pis-*/mie-*, which exist at no release. --- non-finding [context]: Named the mechanism the battery was measuring, while inside it. Reviewing the PR: the names "look like abbreviations of the CSS logical property names (padding-block-start -> pbs), which is a naming convention some other tools use - and exactly the kind of API an AI or a developer coming from another ecosystem might guess Tailwind has. It doesn't." It correctly identified that other tools use the convention (tailwindcss-logical does, as the v2-b twin named outright), correctly identified guessing as the risk, and then landed on the wrong side of it. ============================================================================== RUN tailwindcss--claude-fable-5--v1--2026-08-31 What Claude Fable 5 gets right about tailwindcss — battery v1, tested 2026-08-31 URL: https://stalepriors.com/runs/tailwindcss--claude-fable-5--v1--2026-08-31 JSON: https://stalepriors.com/data/tailwindcss/fable-5.json Library: tailwindcss 4.3.3 (npm), verified 2026-08-31 Model: Claude Fable 5 (claude-fable-5), stated cutoff 2026-01 Version attribution stops at: 4.1.0 (2025-04-01), lag ~9 months Oldest release it could not place: 4.2.0 (2026-02-18) Battery: tailwindcss/v1, 10 tasks, tool uses during test: 0 Tested: 2026-08-31 Summary: Zero findings, and the most consistently version-aware run of the three: every task that had a v3 form volunteered the v3 form explicitly labelled as v3, after giving the v4 answer. Version attribution stops at 4.1.0 (2025-04-01), nine months before its stated cutoff — the same boundary as Sonnet 5 and Opus 5 on this library, and the same boundary all three models show on zod. With zero findings, nothing in the code contradicts its post-4.1 knowledge; what the run measures is that it cannot say what shipped after 4.1.0. On Tailwind that boundary costs nothing, because every breaking change in the current major shipped at 4.0.0, three months before it. --- non-finding [correct] @tailwind base / @tailwind components / @tailwind utilities: Task 1: `@tailwindcss/vite` plugin plus `@import "tailwindcss";`, then spelled out the v3 path it was NOT using — "no tailwind.config.js, no postcss.config.js, no content array, no npx tailwindcss init" — and gave the v3 commands for contrast. This is the behaviour the code-vs-claim rule is meant to reward: v4 leads, v3 is labelled legacy. --- non-finding [correct] tailwind.config.js: Task 2: `@theme` with --color-brand and --font-display, and named the v3 equivalent (theme.extend.colors.brand) as the v3 equivalent. --- non-finding [correct] @layer utilities / @layer components (for defining custom classes): Task 3: `@utility tab-4`. Correctly noted that in v3 an @layer utilities rule did get variant support — which is true of v3 and is exactly the behaviour v4 removed. --- non-finding [correct] safelist / corePlugins / separator: Task 4: `@source inline(...)` including the range form `text-slate-{100..900..100}`, verified against the docs. Named safelist explicitly as the v3 equivalent, and recommended the static lookup map as the better fix. --- non-finding [correct] flex-shrink-* / flex-grow-* / overflow-ellipsis / decoration-slice / decoration-clone: Task 5: `shrink-0`, `flex-1`, `min-w-0`, with the min-width:auto explanation. --- non-finding [correct] bg-opacity-* / text-opacity-* / border-opacity-* / ring-opacity-* / placeholder-opacity-* / divide-opacity-*: Task 6: `bg-black/50`. --- non-finding [correct] bg-[--var] (CSS variable shorthand in arbitrary values): Task 7: `bg-[var(--brand)]` as the primary answer and `bg-(--brand)` named as the v4 shorthand. Both work on 4.x. --- non-finding [correct] border / divide (default colour): Task 8: currentColor border default correctly attributed to v4 with gray-200 as the v3 behaviour, and the shadow scale rename stated in both directions (v4 shadow-sm == v3 shadow; v3 shadow-sm == v4 shadow-xs). --- non-finding [correct] ring: Task 9: bare ring is 1px currentColor in v4, 3px blue-500/50 in v3, plus the correct restoration (`ring-3 ring-blue-500/50`). Also reached for focus-visible rather than focus unprompted. --- non-finding [correct] @apply inside Vue/Svelte/Astro ``` > `@reference "tailwindcss";` is sufficient only if you use no custom theme tokens. Using the generated CSS variables directly (`color: var(--color-brand)`) avoids the problem and is faster. *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Using @apply with Vue, Svelte, or CSS modules"](https://tailwindcss.com/docs/upgrade-guide) #### @layer utilities / @layer components (for defining custom classes) **Removed in tailwindcss 4.0.0** (2025-01-21) v4 uses native CSS cascade layers and no longer hijacks the `@layer` at-rule, so a class defined inside `@layer utilities` is no longer registered as a utility and gets no variant support (`hover:`, `md:`). Custom utilities are declared with the `@utility` API, which also accepts functional forms. *The stale belief:* That wrapping a class in `@layer utilities` is how you add a custom utility that works with variants. ```css /* Stale */ @layer utilities { .tab-4 { tab-size: 4; } } /* Current */ @utility tab-4 { tab-size: 4; } /* functional form, matching tab-2, tab-4, tab-[13] */ @utility tab-* { tab-size: --value(integer); } ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Adding custom utilities"](https://tailwindcss.com/docs/upgrade-guide) · [Tailwind CSS docs — Adding custom styles, functional utilities](https://tailwindcss.com/docs/adding-custom-styles) #### @tailwind base / @tailwind components / @tailwind utilities **Removed in tailwindcss 4.0.0** (2025-01-21) The three `@tailwind` directives were removed. A v4 CSS entry file pulls the framework in with a plain CSS import: `@import "tailwindcss";`. *The stale belief:* That a Tailwind entry stylesheet starts with `@tailwind base; @tailwind components; @tailwind utilities;`. ```css /* Stale */ @tailwind base; @tailwind components; @tailwind utilities; /* Current */ @import "tailwindcss"; ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Removed @tailwind directives"](https://tailwindcss.com/docs/upgrade-guide) #### bg-[--var] (CSS variable shorthand in arbitrary values) **Renamed in tailwindcss 4.0.0** (2025-01-21) The bare-variable shorthand inside arbitrary values moved from square brackets to parentheses: `bg-[--brand]` becomes `bg-(--brand)`. The explicit longhand `bg-[var(--brand)]` is unaffected and still works. *The stale belief:* That `bg-[--brand]` resolves the custom property. ```html
``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Variables in arbitrary values"](https://tailwindcss.com/docs/upgrade-guide) #### `bg-opacity-* / text-opacity-* / border-opacity-* / ring-opacity-* / placeholder-opacity-* / divide-opacity-*` **Removed in tailwindcss 4.0.0** (2025-01-21) The separate opacity utilities were removed. Use the slash opacity modifier on the colour utility itself. *The stale belief:* That translucency is expressed as a second class, `bg-black bg-opacity-50`. ```html
``` > No CSS is emitted for the removed class, so the element is simply opaque. Nothing errors — the failure is visual. *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Removed deprecated utilities"](https://tailwindcss.com/docs/upgrade-guide) #### commas in grid-cols-[…] / grid-rows-[…] / object-[…] arbitrary values **Stricter in tailwindcss 4.0.0** (2025-01-21) v3 replaced commas with spaces inside these particular arbitrary values, a compatibility hack carried over from v2. That is gone: use underscores for spaces, as everywhere else in arbitrary values. *The stale belief:* That `grid-cols-[max-content,auto]` means two columns. ```html
``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Arbitrary values in grid and object-position utilities"](https://tailwindcss.com/docs/upgrade-guide) #### container (center / padding configuration) **Removed in tailwindcss 4.0.0** (2025-01-21) The `container` utility's `center` and `padding` configuration options no longer exist. Extend the utility itself instead. *The stale belief:* That container centring and padding are set under `theme.container` in the JS config. ```css /* Stale */ // tailwind.config.js theme: { container: { center: true, padding: '2rem' } } /* Current */ @utility container { margin-inline: auto; padding-inline: 2rem; } ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Container configuration"](https://tailwindcss.com/docs/upgrade-guide) #### `prefix` **Renamed in tailwindcss 4.0.0** (2025-01-21) Prefixes now look like variants and sit at the front of the whole class: `tw:flex`, `tw:hover:bg-red-600`. The prefix is declared on the import, and theme variables are still authored unprefixed. *The stale belief:* That the prefix is glued to the utility name after any variants, `hover:tw-bg-red-600`. ```html
/* CSS */ @import "tailwindcss" prefix(tw);
``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Using a prefix"](https://tailwindcss.com/docs/upgrade-guide) #### `resolveConfig` **Removed in tailwindcss 4.0.0** (2025-01-21) The `resolveConfig` export was removed. Theme values are real CSS custom properties at runtime, so reference them directly (`var(--color-blue-500)`), or read one in JS with `getComputedStyle(document.documentElement).getPropertyValue("--shadow-xl")`. *The stale belief:* That you import `resolveConfig` from `tailwindcss/resolveConfig` to get theme values into JavaScript. ```js // Stale import resolveConfig from 'tailwindcss/resolveConfig'; import config from './tailwind.config.js'; const { theme } = resolveConfig(config); // Current const styles = getComputedStyle(document.documentElement); const shadow = styles.getPropertyValue("--shadow-xl"); // or just use the variable in CSS/JS values: ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Theme values in JavaScript"](https://tailwindcss.com/docs/upgrade-guide) #### `safelist / corePlugins / separator` **Removed in tailwindcss 4.0.0** (2025-01-21) The `safelist`, `corePlugins` and `separator` options are not supported in v4, from a JS config or anywhere else. Their replacement did **not** ship with the removal: `@source inline(...)` — whose argument is brace-expanded — arrived in **4.1.0 (2025-04-01)**, ten weeks after 4.0.0 removed `safelist`. On 4.0.x there is no safelisting mechanism at all, and `@source inline("...")` is a build error there: the 4.0 parser accepts a quoted path only and throws ``@source` paths must be quoted.` *The stale belief:* That runtime-assembled class names are rescued by a `safelist` array in `tailwind.config.js`. **The code below needs tailwindcss 4.1.0 or later.** Between 4.0.0 and 4.1.0 this correction does not apply — see the note. ```css /* Stale */ // tailwind.config.js module.exports = { safelist: ['bg-red-500', { pattern: /bg-(red|green)-(100|500)/, variants: ['hover'] }], }; /* Current */ @import "tailwindcss"; @source inline("{hover:,}bg-{red,green}-{100,500}"); @source inline("{hover:,}bg-red-{50,{100..900..100},950}"); /* ranges work too */ /* on 4.0.x there is no replacement: write the complete class names into a lookup map so the scanner sees them literally. */ ``` > Bisected in the published packages 2026-09-02: the `@source` parser handles `inline(` (and `not `) from **4.1.0**; the string is absent from the dist of 4.0.0, 4.0.9, 4.0.12, 4.0.15 and 4.0.17 — the last 4.0.x — where the parser takes a quoted path only. The upgrade-guide sentence quoted below is on the current docs page and describes 4.1+, not the 4.0.0 release it appears under; this fact repeated that gap until it was re-verified. The better fix where you control the code is still to write complete class names into a lookup map so the scanner sees them; `@source inline()` is for strings that genuinely cannot be static. *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Using a JavaScript config file"](https://tailwindcss.com/docs/upgrade-guide) · [Tailwind CSS docs — Detecting classes in source files, "Safelisting specific utilities"](https://tailwindcss.com/docs/detecting-classes-in-source-files) · [tailwindcss 4.1.0 — shipped package (dist), @source inline( parser present; absent through 4.0.17](https://github.com/tailwindlabs/tailwindcss/releases/tag/v4.1.0) · 2025-04-01 #### `Sass / Less / Stylus` **New requirement in tailwindcss 4.0.0** (2025-01-21) v4 is not designed to be used with a CSS preprocessor. Tailwind is the preprocessor: you cannot use Sass, Less or Stylus for your stylesheets or for ` ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Default border color"](https://tailwindcss.com/docs/upgrade-guide) #### hover variant **Behaviour changed in tailwindcss 4.0.0** (2025-01-21) `hover:` is now wrapped in `@media (hover: hover)`, so it does not fire on touch devices that previously triggered hover on tap. Treat hover as an enhancement; if you genuinely depend on the old behaviour, override the variant. *The stale belief:* That `hover:` styles apply on touch devices when the user taps. ```css /* Stale */ /* v3 */ .hover\:underline:hover { text-decoration: underline; } /* Current */ /* v4 */ @media (hover: hover) { .hover\:underline:hover { text-decoration: underline; } } /* to opt out: */ @custom-variant hover (&:hover); ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Hover styles on mobile"](https://tailwindcss.com/docs/upgrade-guide) #### `outline-none` **Renamed in tailwindcss 4.0.0** (2025-01-21) In v3, `outline-none` did not set `outline-style: none` — it set an invisible outline that still appeared in forced-colors mode for accessibility. That behaviour is now called `outline-hidden`. The v4 `outline-none` really does remove the outline. Code carrying `focus:outline-none` from v3 therefore loses its forced-colors-mode focus indicator. *The stale belief:* That `focus:outline-none` is the safe way to suppress the browser's focus ring before drawing your own. ```html ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Renamed outline utility"](https://tailwindcss.com/docs/upgrade-guide) #### Preflight base styles **Default changed in tailwindcss 4.0.0** (2025-01-21) Three Preflight defaults changed: placeholder text uses the current text colour at 50% opacity rather than `gray-400`; buttons use `cursor: default` rather than `cursor: pointer`; and margins are reset on ``, so dialogs are no longer centred by default. Each is restorable with a small `@layer base` block. *The stale belief:* That buttons get a pointer cursor and dialogs are centred out of the box. ```css @layer base { button:not(:disabled), [role="button"]:not(:disabled) { cursor: pointer; } input::placeholder, textarea::placeholder { color: var(--color-gray-400); } dialog { margin: auto; } } ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Preflight changes"](https://tailwindcss.com/docs/upgrade-guide) #### `ring` **Default changed in tailwindcss 4.0.0** (2025-01-21) The bare `ring` utility changed from a 3px `blue-500` ring to a 1px `currentColor` ring. A v3 focus style written as `focus:ring` is therefore now a thin ring in the element's own text colour — on a coloured button with white text, effectively invisible. Use `ring-3` plus an explicit `ring-` to restore the v3 appearance. *The stale belief:* That `focus:ring` alone produces the familiar 3px blue focus ring. ```html ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Default ring width and color"](https://tailwindcss.com/docs/upgrade-guide) #### `shadow-sm / shadow / rounded-sm / rounded / blur-sm / blur` **Renamed in tailwindcss 4.0.0** (2025-01-21) The shadow, radius and blur scales were shifted one step so every utility has a named value. What v3 called `shadow` is now `shadow-sm`, and v3's `shadow-sm` is now `shadow-xs`. The same applies to `rounded`/`rounded-sm`, `blur`/`blur-sm` and `drop-shadow`/`backdrop-blur`. The old bare names still resolve, so nothing errors — they just render one step larger than the author intended. *The stale belief:* That `shadow-sm` is the subtlest shadow in the scale. ```html
``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Updated shadow, radius, and blur scales"](https://tailwindcss.com/docs/upgrade-guide) #### space-x-* / space-y-* / divide-x-* / divide-y-* (selector) **Behaviour changed in tailwindcss 4.0.0** (2025-01-21) The selector behind these utilities changed from `> :not([hidden]) ~ :not([hidden])` to `> :not(:last-child)`, for performance, and the spacing now hangs off the *bottom* of each child rather than the top of each sibling. Layouts using these with inline elements, or with per-child margin tweaks, can shift. Prefer `flex`/`grid` with `gap`. *The stale belief:* That `space-y-4` adds top margin to every child after the first. ```css /* Stale */ /* v3 */ .space-y-4 > :not([hidden]) ~ :not([hidden]) { margin-top: 1rem; } /* Current */ /* v4 */ .space-y-4 > :not(:last-child) { margin-bottom: 1rem; } /* recommended instead: */ /*
*/ ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Space-between selector"](https://tailwindcss.com/docs/upgrade-guide) #### stacked variant order **Behaviour changed in tailwindcss 4.0.0** (2025-01-21) Stacked variants now apply left to right instead of right to left, to read more like CSS. Order-sensitive stacks must be reversed: v3's `first:*:pt-0` is v4's `*:first:pt-0`. In practice this bites the direct-child variant `*` and plugin variants like `prose-headings`. *The stale belief:* That `first:*:pt-0` targets the first child. ```html
``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Variant stacking order"](https://tailwindcss.com/docs/upgrade-guide) #### `transform-none / rotate-* / scale-* / translate-*` **Behaviour changed in tailwindcss 4.0.0** (2025-01-21) These utilities are now built on the individual `rotate`, `scale` and `translate` CSS properties rather than the `transform` shorthand. Two consequences: `transform-none` no longer resets them (reset the individual property, e.g. `scale-none`), and a custom transition list containing `transform` no longer transitions them — list the individual properties, `transition-[opacity,scale]`. *The stale belief:* That `transform-none` clears a scale or rotation, and that `transition-[opacity,transform]` animates `scale-*`. ```html ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Individual transform properties"](https://tailwindcss.com/docs/upgrade-guide) ### Deprecated, or a better API now exists Works today. It is the older idiom, and some of it is scheduled for removal. #### `@container-size` **Added in tailwindcss 4.3.0** (2026-05-08) Since 4.3.0 the `@container-size` utility sets `container-type: size`, alongside the long-standing `@container` utility which sets `container-type: inline-size`. Use `@container-size` where a container query needs to read the container's block size as well as its inline size; `@container` is unchanged and is still the right default. *The stale belief:* That `@container` is the only container-type utility, so `container-type: size` needs `[container-type:size]` or a hand-written rule. ```html
``` > S3 under the workaround rule: `[container-type:size]` compiles at 4.2.4 and at 4.3.0 and emits the same declaration. VERIFIED BY COMPILATION 2026-09-06 against installed 4.2.4, 4.3.0 and 4.3.3. `@container-size` emits nothing at 4.2.4 and `container-type: size` at 4.3.0 and 4.3.3; `@container` emits `container-type: inline-size` at all three, so the older utility did not change. *Reproduced against: **Claude Fable 5.1** (S2) — tailwindcss/v3-a, 2026-09-06.* Source: [tailwindcss v4.3.0 release notes — Added](https://github.com/tailwindlabs/tailwindcss/releases/tag/v4.3.0) · 2026-05-08 · [tailwindcss 4.3.0, shipped package — `@container-size` compiles to `container-type: size`; `@container` still compiles to `container-type: inline-size`.](https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.3.0.tgz) · [tailwindcss 4.2.4, shipped package — `@container-size` compiles to nothing; `@container` is already present.](https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.2.4.tgz) #### @variant (stacked and compound variants) **Added in tailwindcss 4.3.0** (2026-05-08) Since 4.3.0 the `@variant` at-rule accepts a stacked variant (`@variant hover:focus { … }`, applying when both conditions hold) and a comma-separated compound list (`@variant hover, focus { … }`, applying when any of them holds). Both spellings work with modifier-carrying variants too, so `@variant group-hover:focus { … }` is legal. Before 4.3.0 `@variant` took exactly one variant name and anything else was a BUILD ERROR: `Cannot use @variant with unknown variant: hover:focus`. *The stale belief:* That `@variant` takes a single variant only, so combining two conditions means nesting one `@variant` block inside another, and a comma-separated list has no equivalent at all — a model holding this belief will tell the author of a working `@variant hover:focus { … }` block that Tailwind cannot parse it. ```css /* Stale */ @import "tailwindcss"; .btn { @variant hover { @variant focus { outline: 2px solid magenta; } } } /* Current */ @import "tailwindcss"; .btn { @variant hover:focus { outline: 2px solid magenta; } /* and the compound form, which nesting cannot express at all */ @variant hover, focus, active { outline-offset: 2px; } } ``` > SEVERITY IS ASYMMETRIC ACROSS THE TWO HALVES AND THE FACT DOES NOT AVERAGE THEM. The stacked half is S3: the nested workaround compiles at 4.2.4 and at 4.3.0 and, verified by compilation, produces BYTE-IDENTICAL output to the stacked form at 4.3.0, so a reader who nests loses nothing. The compound half has no nested equivalent — nesting means AND, the comma means OR — so a reader who believes the comma is illegal writes the block out once per variant instead. That is still working CSS, which is why the fact as a whole stays S3 rather than S2; the S2 direction is review, where a draw tells the author of a working `@variant hover:focus` block that it is a build error. VERIFIED BY COMPILATION 2026-09-06 with the real engine. At 4.2.4 `@variant hover:focus`, `@variant group-hover:focus` and `@variant hover, focus, active` each THROW `Cannot use @variant with unknown variant: …` — a hard build failure, not a silent no-op — while `@variant hover` and the nested form compile. At 4.3.0 and 4.3.3 all five compile. *Reproduced against: **Claude Fable 5.1** (S3) — tailwindcss/v3-a, 2026-09-06.* Source: [tailwindcss v4.3.0 release notes — Added](https://github.com/tailwindlabs/tailwindcss/releases/tag/v4.3.0) · 2026-05-08 · [tailwindcss v4.3.0 release notes — Added](https://github.com/tailwindlabs/tailwindcss/releases/tag/v4.3.0) · 2026-05-08 · [tailwindcss 4.3.0, shipped package — `@variant hover:focus { outline: 2px solid magenta }` compiles to `&:hover { @media (hover: hover) { &:focus { outline: 2px solid magenta } } }`, identical to the nested form; `@variant hover, focus { … }` compiles to two separate rules.](https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.3.0.tgz) · [tailwindcss 4.2.4, shipped package — the same stacked and compound sources throw `Cannot use @variant with unknown variant: hover:focus` / `: hover, focus`; the single-variant and nested forms compile.](https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.2.4.tgz) #### `scrollbar-auto / scrollbar-thin / scrollbar-none / scrollbar-thumb-* / scrollbar-track-* / scrollbar-gutter-*` **Added in tailwindcss 4.3.0** (2026-05-08) Since 4.3.0 Tailwind ships first-class scrollbar utilities and no plugin is needed. `scrollbar-auto`, `scrollbar-thin` and `scrollbar-none` set `scrollbar-width`. `scrollbar-thumb-*` and `scrollbar-track-*` set `scrollbar-color` — each writes one half into a CSS variable (`--tw-scrollbar-thumb` / `--tw-scrollbar-track`) and both emit the same `scrollbar-color` declaration, so use them as a pair or the unset half falls back to the variable's default. `scrollbar-gutter-auto` and `scrollbar-gutter-stable` set `scrollbar-gutter`. Colour utilities take the theme palette and arbitrary values alike: `scrollbar-thumb-slate-500` and `scrollbar-thumb-[#123456]` both compile. Do not install `tailwind-scrollbar` and do not hand-write a `::-webkit-scrollbar` block for the standard properties. *The stale belief:* That Tailwind has no scrollbar utilities, so styling a scrollbar means installing the third-party `tailwind-scrollbar` plugin, writing `::-webkit-scrollbar` rules in your own CSS, or reaching for arbitrary properties. A model holding this belief will also reject working `scrollbar-thin scrollbar-thumb-slate-500` markup in review as classes that generate nothing. ```html
``` > SEVERITY IS S3 AND THE REASON MATTERS, exactly as on LF30. The arbitrary-property workaround a stale model offers (`[scrollbar-width:thin]`, `[scrollbar-color:red_blue]`) compiles at 4.2.4 AND at 4.3.0 and emits the identical declaration, so acting on the stale belief costs verbosity, not correctness, and the workaround rule (HARNESS.md, JOURNAL/030) forbids charging it as a defect. Two directions are worse and are S2 where a run shows them: telling the author of working `scrollbar-thin` markup that it generates nothing, and sending the reader to a third-party plugin or to `::-webkit-scrollbar` — the latter is not equivalent, because `::-webkit-scrollbar` is unimplemented in Firefox while `scrollbar-width` and `scrollbar-color` are the standard properties this family emits. VERIFIED BY COMPILATION 2026-09-06: every class in `correct_code` built one at a time with the real engine against installed 4.1.0, 4.2.4, 4.3.0 and 4.3.3. All of them emit nothing at 4.1.0 and 4.2.4 and the expected declaration at 4.3.0 and 4.3.3 — a clean bisection onto the target release, with the whole 4.2.x line ruled out rather than just 4.2.0. THE SCHEME DOES NOT GENERALISE, and this is the family's guessing trap: `scrollbar-corner-*`, `scrollbar-hidden`, `scrollbar-width-thin` and `scrollbar-thumb-rounded` are absent from 4.2.4, 4.3.0 and 4.3.3 alike and compile to nothing at all three. There is no corner utility; CSS `scrollbar-color` takes exactly two colours and the utilities cover exactly those two. *Reproduced against: **Claude Fable 5.1** (S2) — tailwindcss/v3-a, 2026-09-06.* Source: [tailwindcss v4.3.0 release notes — Added](https://github.com/tailwindlabs/tailwindcss/releases/tag/v4.3.0) · 2026-05-08 · [tailwindcss v4.3.0 release notes — Added](https://github.com/tailwindlabs/tailwindcss/releases/tag/v4.3.0) · 2026-05-08 · [tailwindcss 4.3.0, shipped package — the engine compiles `scrollbar-thin` to `scrollbar-width: thin`, `scrollbar-thumb-red-500` to `--tw-scrollbar-thumb: var(--color-red-500); scrollbar-color: var(--tw-scrollbar-thumb) var(--tw-scrollbar-track)`, and `scrollbar-gutter-stable` to `scrollbar-gutter: stable`.](https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.3.0.tgz) · [tailwindcss 4.2.4, shipped package — every one of those class names compiles to nothing, at the last release of the 4.2 line. The family begins at 4.3.0, not earlier in 4.2.x.](https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.2.4.tgz) · [tailwindcss 4.3.3, shipped package — `scrollbar-corner-red-500`, `scrollbar-hidden`, `scrollbar-width-thin` and `scrollbar-thumb-rounded` are absent and compile to nothing at the newest release that exists.](https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.3.3.tgz) · [Tailwind CSS docs — scrollbar-width (Interactivity), documented at v4.3](https://tailwindcss.com/docs/scrollbar-width) #### `zoom-* / tab-*` **Added in tailwindcss 4.3.0** (2026-05-08) Since 4.3.0 Tailwind ships `zoom-*` utilities for the CSS `zoom` property and `tab-*` utilities for `tab-size`. `zoom-*` takes a percentage-style number — `zoom-50` emits `zoom: 50%`, `zoom-150` emits `zoom: 150%` — and `tab-*` takes a plain integer, so `tab-4` emits `tab-size: 4`. Both accept arbitrary values: `zoom-[1.5]` emits `zoom: 1.5` (a bare ratio, not a percentage) and `tab-[3]` emits `tab-size: 3`. *The stale belief:* That neither CSS `zoom` nor `tab-size` has a Tailwind utility, so both need arbitrary properties or hand-written CSS. ```html


``` > S3 for the same reason as LF31: `[tab-size:4]` and `[zoom:1.5]` compile at 4.2.4 and at 4.3.0 and emit the identical declaration, so the stale belief costs verbosity rather than correctness and cannot be charged as a defect under the workaround rule. Note the unit asymmetry, which is the one place a reader can go wrong translating between the two spellings: the named scale is a percentage (`zoom-50` -> `zoom: 50%`) while the arbitrary form is passed through untouched (`zoom-[1.5]` -> `zoom: 1.5`, which is 150%, not 1.5%). `zoom-normal` and `zoom-reset` do not exist. VERIFIED BY COMPILATION 2026-09-06 against installed 4.1.0, 4.2.4, 4.3.0 and 4.3.3: `zoom-50`, `zoom-25`, `zoom-150`, `tab-1`, `tab-4` and `tab-8` emit nothing at 4.1.0 and 4.2.4 and the expected declaration at 4.3.0 and 4.3.3. *Reproduced against: **Claude Fable 5.1** (S2) — tailwindcss/v3-a, 2026-09-06.* Source: [tailwindcss v4.3.0 release notes — Added](https://github.com/tailwindlabs/tailwindcss/releases/tag/v4.3.0) · 2026-05-08 · [tailwindcss v4.3.0 release notes — Added](https://github.com/tailwindlabs/tailwindcss/releases/tag/v4.3.0) · 2026-05-08 · [tailwindcss 4.3.0, shipped package — `zoom-50` compiles to `zoom: 50%`, `zoom-[1.5]` to `zoom: 1.5`, `tab-4` to `tab-size: 4`.](https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.3.0.tgz) · [tailwindcss 4.2.4, shipped package — `zoom-50`, `zoom-[1.5]`, `tab-4` all compile to nothing, while `[zoom:1.5]` and `[tab-size:4]` compile there and at 4.3.0 alike.](https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.2.4.tgz) · [Tailwind CSS docs — zoom, documented at v4.3](https://tailwindcss.com/docs/zoom) · [Tailwind CSS docs — tab-size, documented at v4.3](https://tailwindcss.com/docs/tab-size) #### `pbs-* / pbe-* / mbs-* / mbe-* / border-bs-* / inline-* / block-* / min-inline-* / max-block-* / inset-bs-*` **Added in tailwindcss 4.2.0** (2026-02-18) Since 4.2.0 Tailwind ships first-class utilities for the CSS block axis and for logical sizing. Spacing: `pbs-*`, `pbe-*`, `mbs-*`, `mbe-*`, and the scroll variants `scroll-pbs-*`, `scroll-pbe-*`, `scroll-mbs-*`, `scroll-mbe-*`. Borders: `border-bs-*`, `border-be-*`. Sizing: `inline-*`, `min-inline-*`, `max-inline-*` for `inline-size` and its bounds, and `block-*`, `min-block-*`, `max-block-*` for `block-size`. Insets: `inset-s-*`, `inset-e-*`, `inset-bs-*`, `inset-be-*`. Do not reach for `[padding-block-start:...]` arbitrary values and do not substitute the physical `pt-*` / `mb-*` / `h-*` utilities, which are not equivalent in a vertical writing mode. The INLINE axis is the older half of the family and keeps its v3 spelling: `ps-*`, `pe-*`, `ms-*`, `me-*`. There is no `pis-*`, `pie-*`, `mis-*`, `mie-*` or `border-is-*` at any release, 4.3.3 included. *The stale belief:* That Tailwind covers the inline axis only (`ps-*`, `ms-*`, `start-*`, `end-*`) and offers nothing for `padding-block-start`, `border-block-start`, `inline-size` or `block-size`, so a writing-mode-agnostic component needs arbitrary values or hand-written CSS. A model holding this belief will also reject working `pbs-6` / `inline-full` markup in review as classes that generate nothing. ```html
``` > SEVERITY IS S3 AND THE REASON MATTERS. The arbitrary-value workaround a stale model offers (`[padding-block-start:1.5rem]`) compiles and emits the same declaration, so acting on the stale belief costs verbosity, not correctness - the workaround rule (HARNESS.md, JOURNAL/030) forbids charging that as a defect. The S2 direction is REVIEW: a model that tells the author of working `pbs-6` markup that the class does not exist and emits nothing is asserting something the compiler contradicts, and a finding charged in that direction states S2 on the run. Substituting `pt-*` for `pbs-*` is also S2, but only in a non-horizontal writing mode - in `writing-mode: horizontal-tb` the two are identical, which is why this fact does not claim the substitution is broken in general. VERIFIED BY COMPILATION 2026-09-03, every class in `correct_code` built one at a time against installed 4.1.0, 4.2.0 and 4.3.3 with the real engine. Every 4.2.0 name emits nothing at 4.1.0 and the expected declaration at 4.2.0; `ps-4` and `me-4` emit at both; `pis-4`, `pie-4`, `mis-4`, `mie-4` and `border-is-2` emit at none of the three. That last check is the point of the `pis-*` sentence in the statement: the family has a naming scheme a reader (or a model) will over-generalise, and four of the names it suggests do not exist. Not to be confused with the display utilities. `inline`, `block`, `inline-block` and `inline-flex` are unchanged and still set `display`; `inline-4` and `block-full` are the new sizing utilities and do not collide with them. *Reproduced against: **Claude Opus 5** (S2) — tailwindcss/v2-a, 2026-09-03.* Source: [tailwindcss CHANGELOG - 4.2.0 (2026-02-18), Added](https://raw.githubusercontent.com/tailwindlabs/tailwindcss/main/CHANGELOG.md) · 2026-02-18 · [tailwindcss CHANGELOG - 4.2.0 (2026-02-18), Added](https://raw.githubusercontent.com/tailwindlabs/tailwindcss/main/CHANGELOG.md) · 2026-02-18 · [tailwindcss CHANGELOG - 4.2.0 (2026-02-18), Added](https://raw.githubusercontent.com/tailwindlabs/tailwindcss/main/CHANGELOG.md) · 2026-02-18 · [tailwindcss 4.2.0, shipped package - dist/ registers pbs, pbe, mbs, mbe, border-bs, border-be, inset-s, inset-e, inset-bs, inset-be, scroll-pbs and scroll-mbe, and the engine compiles `pbs-6` to `padding-block-start: calc(var(--spacing) * 6)` and `inline-full` to `inline-size: 100%`.](https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.2.0.tgz) · [tailwindcss 4.1.0, shipped package - none of those names is registered and every one of them compiles to nothing, while `ps-4` and `me-4` compile normally. The family begins at 4.2.0 and the inline axis predates it.](https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.1.0.tgz) · [tailwindcss 4.3.3, shipped package - `pis-*`, `pie-*`, `mis-*`, `mie-*` and `border-is-*` are absent from dist/ and compile to nothing at the newest release that exists. The scheme does not generalise to the inline axis.](https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.3.3.tgz) #### `start-* / end-*` **Deprecated in tailwindcss 4.2.0** (2026-02-18) As of 4.2.0 the logical inset utilities `start-*` and `end-*` are deprecated in favour of `inset-s-*` and `inset-e-*`, alongside a new family of logical-property utilities (`pbs-*`, `mbe-*`, `border-bs-*`, `inline-*`, `block-*`). *The stale belief:* That `start-0` / `end-0` are the current logical inset utilities. ```html
``` > Published 2026-02-18, after the stated cutoff of two of the three models in this Index. Not chargeable against them - a scheduled retest, not a pass. VERIFIED AGAINST THE COMPILER 2026-09-03, because JOURNAL/041 had just found two tailwindcss facts that read a status word out of a release note and were wrong about the artifact. This one is right, and deprecated really does mean deprecated: at 4.2.0 `start-0` and `inset-s-0` compile to the SAME declaration (`inset-inline-start: calc(var(--spacing) * 0)`), so nobody's working markup broke. Do not restate this fact as `removed`. The only difference the compiler shows is cosmetic and appears later: at 4.3.3 `inset-s-0` emits the constant-folded `inset-inline-start: 0px` while `start-0` still emits the `calc()` form. Same computed value. *Reproduced against: **Claude Opus 5** (S3) — tailwindcss/v2-a, 2026-09-03.* Source: [tailwindcss CHANGELOG — 4.2.0 (2026-02-18), Deprecated](https://raw.githubusercontent.com/tailwindlabs/tailwindcss/main/CHANGELOG.md) · 2026-02-18 · [tailwindcss 4.2.0, shipped package - both spellings compile, and to the same declaration. `start-0` -> `inset-inline-start: calc(var(--spacing) * 0)`; `inset-s-0` -> the same string.](https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.2.0.tgz) · [tailwindcss 4.1.0, shipped package - `inset-s-0` and `inset-e-0` generate no CSS at all, and dist/ registers only the `start` and `end` utility names. The deprecation and its replacement both begin at 4.2.0.](https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.1.0.tgz) #### `bg-gradient-*` **Deprecated in tailwindcss 4.0.0** (2025-01-21) Linear gradient utilities were renamed from `bg-gradient-*` to `bg-linear-*`, making room for the new `bg-radial-*` and `bg-conic-*` families and for angle values like `bg-linear-45`. The old spelling was kept as an alias, not removed: 4.x rewrites a `bg-gradient-to-*` candidate to `bg-linear-to-*` before compiling it, so v3 markup still renders a gradient. Adopt the new name for the angle and radial/conic families, not because the old one stops working. Gradients also keep their other stops when a variant overrides one — use `via-none` to drop back to a two-stop gradient in a given state. *The stale belief:* That `bg-gradient-to-r` is the current spelling of the linear gradient utility. It is the v3 spelling, retained as an alias. ```html
``` > VERIFIED 2026-09-03 (JOURNAL/041), against the shipped packages, and the statement was sharpened as a result. It previously said only that the utilities were "renamed", which reads as the old name being gone; the note under it hedged that the old name had not been verified to stop working. It has now been verified to keep working. `tools/audit/css-audit.mjs` built `bg-gradient-to-r` against tailwindcss 4.0.0 and 4.3.3: it generates at both, and at 4.0.0 its declarations are identical to `bg-linear-to-r`'s. The alias is explicit in the shipped bundle — the candidate parser rewrites a root beginning `bg-gradient-to-` to `bg-linear-to-` — which is why the two directional families cannot diverge. Only `bg-linear-45` and the `bg-radial-*` / `bg-conic-*` families are genuinely new surface. *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS v4.0 announcement — Expanded gradient APIs](https://tailwindcss.com/blog/tailwindcss-v4) · 2025-01-22 · [tailwindcss CHANGELOG — bg-linear-* alias](https://raw.githubusercontent.com/tailwindlabs/tailwindcss/main/CHANGELOG.md) · [tailwindcss 4.3.3, shipped package — dist/lib.mjs rewrites a bg-gradient-to-* root to bg-linear-to-*](https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.3.3.tgz) #### `flex-shrink-* / flex-grow-* / overflow-ellipsis / decoration-slice / decoration-clone` **Deprecated in tailwindcss 4.0.0** (2025-01-21) These v3 aliases were dropped from the documentation, not from the compiler. `flex-shrink-*`, `flex-grow-*`, `overflow-ellipsis`, `decoration-slice` and `decoration-clone` are still registered utilities in every 4.x release and emit exactly the same declarations as `shrink-*`, `grow-*`, `text-ellipsis`, `box-decoration-slice` and `box-decoration-clone`. Markup carrying the old names keeps working and renders identically. Adopt the new spellings because the old ones are undocumented and no longer suggested by tooling — not because anything breaks. *The stale belief:* That `flex-shrink-0` is the current, documented spelling. It is the v3 spelling: it still compiles to `flex-shrink: 0`, and it is no longer in the docs. ```html
``` > CORRECTED 2026-09-03 (JOURNAL/041), against the shipped packages. This entry previously filed the five aliases as `removed` at S1 breaks-build, on the upgrade guide's line "We've removed any utilities that were deprecated in v3". That line is about the documentation. The compiler never dropped them: `tools/audit/css-audit.mjs` built each stale class against tailwindcss 4.0.0, 4.1.0, 4.2.0 and 4.3.3 and every one generated CSS byte-identical to the replacement this fact prescribes — `.flex-shrink-0 { flex-shrink: 0; }` beside `.shrink-0 { flex-shrink: 0; }`, and the same for the other four. The mechanism is visible in the shipped bundle: each name is registered with `utilities.functional(...)` or `utilities.static(...)` and paired with `utilities.suggest(, () => [])`, an empty suggestion list — registered, and hidden from tooling. A reader following the old entry was told their working markup broke the build. It does not. Same defect class as valibot LF3 (JOURNAL/040): a release note that is accurate about the intent and wrong about the artifact. *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Removed deprecated utilities" (accurate about the docs, not about the compiler)](https://tailwindcss.com/docs/upgrade-guide) · [tailwindcss 4.0.0, shipped package — dist/lib.mjs still registers flex-shrink, flex-grow, overflow-ellipsis, decoration-slice and decoration-clone](https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.0.0.tgz) · [tailwindcss 4.3.3, shipped package — the same five, each beside an empty utilities.suggest() list](https://registry.npmjs.org/tailwindcss/-/tailwindcss-4.3.3.tgz) #### the important modifier (!flex) **Deprecated in tailwindcss 4.0.0** (2025-01-21) The `!` that marks a utility important moved from the start of the class name to the end: `!flex` becomes `flex!`, `hover:!bg-red-600` becomes `hover:bg-red-600!`. The leading form is still supported for compatibility but is deprecated. *The stale belief:* That the important marker goes after the variants and before the utility. ```html
``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "The important modifier"](https://tailwindcss.com/docs/upgrade-guide) #### theme() function **Deprecated in tailwindcss 4.0.0** (2025-01-21) Prefer the generated CSS variables over `theme()`. The dot path is **deprecated, not removed**: `theme(colors.red.500)` and `theme(screens.xl)` still resolve in v4 and compile to exactly the same declarations as `theme(--color-red-500)` and `theme(--breakpoint-xl)`. Write the CSS-variable form in new code — `var(--color-red-500)` outside media queries, `theme(--breakpoint-xl)` inside them, where CSS variables are not supported — but existing dot-path stylesheets are not broken and do not need an urgent migration. *The stale belief:* That theme values are addressed with dot notation, `theme(colors.red.500)`, and that this is the current idiom rather than a deprecated one. ```css /* Stale */ .my-class { background-color: theme(colors.red.500); } @media (width >= theme(screens.xl)) { } /* Current */ .my-class { background-color: var(--color-red-500); } @media (width >= theme(--breakpoint-xl)) { } ``` > Do not tell readers the dot path stopped working. Verified by compiling the four forms with the installed Tailwind engine at both 4.0.0 and 4.3.3: `theme(colors.red.500)` and `theme(--color-red-500)` both emit `color: oklch(0.637 0.237 25.331)`, and `theme(screens.xl)` and `theme(--breakpoint-xl)` both emit `@media (min-width: 80rem)`. The upgrade guide says you *should* use the variable name; it does not say the dot path is rejected, and the v4 line has gone on fixing dot-path resolution in JS plugins and config files (4.3.3 fixes `theme('colors.foo')` lookups). Fourth fact in this pack found asserting a break that the shipped artifact does not perform — see LF5 and LF27. *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Using the theme() function" (advice, not a removal)](https://tailwindcss.com/docs/upgrade-guide) · [tailwindcss CHANGELOG — 4.3.3 still fixing dot-path theme() resolution](https://github.com/tailwindlabs/tailwindcss/blob/main/CHANGELOG.md) · 2026-07-16 ### Wrong facts about the library Not code — versions, minimums and metadata that models state confidently and get wrong. #### browser support baseline **New requirement in tailwindcss 4.0.0** (2025-01-21) v4 targets Safari 16.4+, Chrome 111+ and Firefox 128+, because it depends on `@property` and `color-mix()` for core features. It will not work in older browsers; projects that must support them should stay on v3.4. *The stale belief:* That upgrading to the current major is browser-support-neutral. *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Tailwind CSS docs — Upgrade guide, "Browser requirements"](https://tailwindcss.com/docs/upgrade-guide) ## Not corrections — recorded for honesty Claims seen in a run but not yet verified against a primary source. Never treated as findings: - The battery's two most derivable-looking probes were derived by the subject furthest below the floor, not the nearest. Claude Sonnet 5 (2026-01) produced the correct `@variant` stacked and compound syntax and the correct `tab-*` scale; Claude Opus 5 (2026-05, three weeks nearer the release) got both wrong and derived only `@container-size`. Nothing in the Index predicts which control derives what, and distance below the floor evidently does not. *(open since 2026-09-06)* - Does the block-axis denial survive a probe that does not mention writing modes? Every task in this battery framed the need through vertical-rl, which is a rare enough layout that a subject may be reasoning "Tailwind would not bother" rather than recalling. A probe that asks for the same utilities in a plain horizontal document would separate the two, and it bounds how far F1 generalises: a belief held only under an exotic framing is a weaker prior than one held in ordinary markup. *(open since 2026-09-03)* - The model gave the v4 shadow-sm value as `0 1px 3px 0 rgb(0 0 0 / 0.1), 0 1px 2px -1px rgb(0 0 0 / 0.1)` and shadow-xs as `0 1px 2px 0 rgb(0 0 0 / 0.05)`. The direction of the rename is verified against the upgrade guide; the exact pixel values were not checked against the shipped theme.css. *(open since 2026-08-31)* - The model gave shadow-sm as `0 1px 3px 0 rgb(0 0 0 / 0.1), 0 1px 2px -1px rgb(0 0 0 / 0.1)`. The direction of the scale rename is verified against the upgrade guide, but the exact pixel values of the v4 shadow scale were not independently checked against the shipped theme.css. *(open since 2026-08-31)* --- *Findings, code and citations: `data/tailwindcss/` — one JSON file and one write-up per model, each finding carrying the release that broke the belief, its publication date and a verbatim quote from the primary source. Corrections: `data/tailwindcss/facts.json`. This file is generated by `tools/build-corrections.mjs`; if the prose and the data ever disagree, that is a bug in the generator, not a stale pack.* ============================================================================== CORRECTION PACK: valibot (https://stalepriors.com/corrections/valibot.md) # valibot correction pack · for projects on valibot@^1.2 ## What we actually measured 5 Claude models were asked for idiomatic valibot code with no tools, purely from training knowledge. 13 reproduced failures across 20 runs, each verified against the release that broke the belief. | Model | Stated cutoff | valibot version attribution stops | Lag inside the window | |---|---|---|---| | Claude Fable 5.1 | 2026-06 | 1.1.0 · 2025-05-06 | ~13 months | | Claude Opus 5 | 2026-05 | 1.1.0 · 2025-05-06 | ~12 months | | Claude Sonnet 5 | 2026-01 | 1.0.0 · 2025-03-19 | ~10 months | | Claude Fable 5 | 2026-01 | 1.1.0 · 2025-05-06 | ~8 months | A recent cutoff is not a defence. Every model here loses track of this library's release history well before the date it states as its own cutoff, and the stopping points cluster far tighter than the cutoffs do. Read that column precisely. It is the newest valibot release whose contents the model can correctly **attribute to that release** — not the newest valibot feature it knows. Past that point a model will often write working code with a newer API while naming the wrong release for it, and that guess runs *early* — it names a release older than the one that shipped the feature. So the question this pack answers is not "does the model know this API" but "can it be trusted about which version the API arrived in" — which is the question that matters when you are pinned to a version. ## How to read an entry Every entry ends with a *Reproduced against* line. Where it names models, we have the generated code that got it wrong, dated, with the model's own words in the run write-up. Where it says no model yet, the correction is verified from the release notes but nothing has been probed for it — it is a fix, not a measurement, and the pack says so rather than blurring the two. The section an entry sits in is the worst case if you act on the stale belief. The severity in brackets after a model's name is what that particular model's output actually did, which can be milder — a model can hold the wrong belief and still, on the day, write code that runs. ## The corrections ### Breaks the build, or throws at runtime Act on the stale belief here and the code does not run. Fix these first. #### `NanoIDAction / NanoIDIssue` **Renamed in valibot 1.1.0** (2025-05-06) The `NanoIDAction` and `NanoIDIssue` interfaces were renamed to `NanoIdAction` and `NanoIdIssue` in 1.1.0. The old casing is not exported and a type import of it fails to compile. *The stale belief:* That the nano-ID action's types spell the acronym in full caps, as they did up to 1.0.0. ```ts // Stale import type { NanoIDAction } from 'valibot' function describe(action: NanoIDAction) { return action.type } // Current import type { NanoIdAction } from 'valibot' function describe(action: NanoIdAction) { return action.type } ``` > A type-level break only: the runtime `nanoid()` action itself was not renamed. It surfaces at compile time, not at runtime. *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [valibot v1.1.0 release notes](https://github.com/open-circle/valibot/releases/tag/v1.1.0) · 2025-05-06 ### Runs, but is silently wrong Nothing errors. The behaviour is simply not what a model trained earlier will tell you. #### `toNumber / toBoolean / toDate / toBigint / toString` **Added in valibot 1.2.0** (2025-11-24) valibot ships dedicated string-to-primitive transformation actions since 1.2.0: `toBigint`, `toBoolean`, `toDate`, `toNumber` and `toString`. Use `toNumber` instead of a bare `transform(Number)`: `toNumber` adds a validation issue when the conversion yields `NaN`, whereas `pipe(string(), transform(Number))` returns `NaN` with `success: true` and hands silently-wrong data downstream. **`toBoolean` is the exception and is not a string parser** — it applies JavaScript `Boolean()`, so the string `"false"` converts to `true`. For the common case of a `"true"`/`"false"` query parameter use `parseBoolean` (1.3.0) or validate the literals with `picklist` before transforming. *The stale belief:* That valibot has no built-in string-to-primitive conversion at all and every such conversion must go through a hand-written `transform(Number)` or `transform((s) => s === "true")`. **The code below needs valibot 1.3.0 or later.** Between 1.2.0 and 1.3.0 this correction does not apply — see the note. ```ts // Stale import * as v from 'valibot' const Query = v.object({ // Number("abc") is NaN, and this pipe returns it with success: true page: v.pipe(v.string(), v.transform(Number)), }) // Current import * as v from 'valibot' const Query = v.object({ // toNumber raises a validation issue instead of yielding NaN page: v.pipe(v.string(), v.toNumber()), // NOT v.toBoolean() — that is Boolean(), and Boolean("false") is true. // parseBoolean (1.3.0) reads the literal words: active: v.pipe(v.string(), v.parseBoolean()), // on valibot < 1.3.0, validate the literals first: // active: v.pipe(v.picklist(["true", "false"]), v.transform((s) => s === "true")), }) ``` > CORRECTED 2026-09-02 (JOURNAL/035), against the shipped package. This entry previously told readers to write `v.pipe(v.string(), v.toBoolean())` for a `"true"`/`"false"` query parameter, and justified the whole fact on the claim that the built-in avoids the `Boolean("false") === true` trap that `transform(Boolean)` falls into. **That justification was false and the recommended code was a bug.** valibot 1.4.2's `toBoolean` is `dataset.value = Boolean(dataset.value)` — literally the same call — so it maps `"false"`, `"0"`, `"no"` and every other non-empty string to `true`. Executed and confirmed: `parse(pipe(string(), toBoolean()), "false")` returns `true`. The word-reading action is `parseBoolean`, which arrived in 1.3.0 (see LF11) and returns `false` for `"false"`. The fact survives on the `toNumber` half, where the built-in genuinely is safer than the hand-roll, and that is now the stated mechanism. The error was caught by the battery `valibot/v2` pre-registration, which required every scored claim to be re-executed against the installed package — and by six test draws that denied these actions exist while independently warning that `Boolean("false")` is `true`. They were right about that and the Index was wrong. ARCHAEOLOGY, 2026-09-02: a subject asserted, inside a run, that valibot USED to have a `coerce` action and that it was deliberately removed. That history is correct. The v0.31.0 migration guide states that `coerce` was removed and why, and maps it onto `pipe` + `unknown` + `transform`. So the belief that produced the wrong answer was not a false memory of the removal but the assumption that nothing replaced it: the generic `coerce` is gone, and the dedicated actions above arrived in 1.2.0, six releases later. *Reproduced against: **Claude Fable 5** (S2), **Claude Opus 5** (S2), **Claude Sonnet 5** (S2) — valibot/v2-a, 2026-09-02; valibot/v2-c, 2026-09-02; valibot/v2-e, 2026-09-02.* Source: [valibot v1.2.0 release notes](https://github.com/open-circle/valibot/releases/tag/v1.2.0) · 2025-11-24 · [valibot 1.4.2, shipped package — the toBoolean implementation](https://registry.npmjs.org/valibot/-/valibot-1.4.2.tgz) · 2026-06-28 · [valibot — Migrate to v0.31.0, Coerce method](https://valibot.dev/guides/migrate-to-v0.31.0/) #### `emoji` **Behaviour changed in valibot 1.2.0** (2025-11-24) The regular expression behind the `emoji` validation action carried a ReDoS vulnerability that was fixed in 1.2.0. If you validate user-supplied text with the `emoji` action, require valibot >= 1.2.0. Advice that recommends the action without that version floor exposes the caller to a denial-of-service on attacker-controlled input. *The stale belief:* That the `emoji` action is safe to run on untrusted, high-volume input at any 1.x version. ```ts // package.json // "valibot": "^1.2.0" <- 1.2.0 fixed the ReDoS in EMOJI_REGEX import * as v from 'valibot' const Reaction = v.pipe(v.string(), v.emoji()) ``` > Verified only to the extent the vendor's own release notes state it: the notes list the fix. This entry claims the fix landed in 1.2.0 and that earlier 1.x is affected; it does not characterise the exploit, which was not published in the release notes. *Reproduced against: **Claude Fable 5** (S2) — valibot/v1, 2026-09-01.* Source: [valibot v1.2.0 release notes](https://github.com/open-circle/valibot/releases/tag/v1.2.0) · 2025-11-24 ### Deprecated, or a better API now exists Works today. It is the older idiom, and some of it is scheduled for removal. #### `guard / parseBoolean / isrc / domain / jwsCompact / cache` **Added in valibot 1.3.0** (2026-03-17) valibot 1.3.0 added the `guard` transformation action (narrow types with a type predicate), `parseBoolean`, the `isrc`, `domain` and `jwsCompact` validation actions, and a `cache` method that caches schema output by input. It also fixed `creditCard` to accept 13-digit Visa numbers. *The stale belief:* That domain-name, ISRC and JWS-compact validation, and schema-level output caching, have no built-in support in valibot. ```ts import * as v from 'valibot' const Host = v.pipe(v.string(), v.domain()) const Token = v.pipe(v.string(), v.jwsCompact()) ``` > PROBED 2026-09-05 by battery valibot/v3 (JOURNAL/055) on the guard and cache halves: all five arms denied both capabilities. Charged twice against Claude Opus 5 on v3-c (F1 guard S3, F2 cache S3). 1.3.0 is admissible against Claude Opus 5 (stated 2026-05) and against Claude Fable 5.1 (stated 2026-06), and below the floor for Claude Sonnet 5 and Claude Fable 5. Both Fable 5.1 arms failed it and neither could be charged: see JOURNAL/055 on the direct-question wording that barred them. The isrc, domain and jwsCompact surfaces were deliberately NOT probed - each is named for the standard it validates, so the probe would measure naming rather than knowledge. CROSS-REFERENCE 2026-09-02 (JOURNAL/035): `parseBoolean` is the action that reads the literal words `"true"`/`"false"` (and `1`/`0`, `yes`/`no`, `on`/`off`, `enabled`/`disabled`, case-insensitively, rejecting anything else). It is **not** interchangeable with 1.2.0's `toBoolean`, which is plain `Boolean()` and maps `"false"` to `true` — see the correction on LF1. A project on valibot < 1.3.0 that needs to read a boolean query parameter has no built-in for it and must validate the literals with `picklist` first. Executed against valibot 1.4.2: `parse(pipe(string(), parseBoolean()), "false")` returns `false`; the empty string is rejected with an issue rather than coerced. *Reproduced against: **Claude Fable 5.1** (S3), **Claude Opus 5** (S3) — valibot/v3-c, 2026-09-05; valibot/v4-a, 2026-09-06.* Source: [valibot v1.3.0 release notes](https://github.com/open-circle/valibot/releases/tag/v1.3.0) · 2026-03-17 · [valibot v1.3.0 release notes — creditCard fix](https://github.com/open-circle/valibot/releases/tag/v1.3.0) · 2026-03-17 · [valibot 1.4.2, shipped package — parseBoolean truthy/falsy word lists](https://registry.npmjs.org/valibot/-/valibot-1.4.2.tgz) · 2026-06-28 #### `toCamelCase / toKebabCase / toPascalCase / toSnakeCase` **Added in valibot 1.4.0** (2026-05-05) valibot 1.4.0 added case-conversion transformation actions (`toCamelCase`, `toKebabCase`, `toPascalCase`, `toSnakeCase`) and the `isoDateTimeSecond` validation action. It also made `intersect` stop mutating its input, so frozen objects and arrays can be merged. *The stale belief:* That valibot has no built-in naming-convention conversions, and that `intersect` cannot be used with frozen inputs. ```ts import * as v from 'valibot' const Slug = v.pipe(v.string(), v.toKebabCase()) ``` > PROBED 2026-09-05 by battery valibot/v3 (JOURNAL/055) on the case-conversion half, in both directions. All five arms denied the capability in the offer direction, and all five rejected a pull request using toKebabCase in the recognition direction - calling a real action an invention. NOTHING CHARGED: 1.4.0 published 2026-05-05, the same month as Claude Opus 5s stated cutoff, so the Index parks it against that subject; Claude Fable 5.1 (stated 2026-06) is the one subject it is admissible against, and both of its arms declined their stated cutoff. Five reproduced failures, zero charged, every one flagged chargeable_miss. Verified this session by installing 1.0.0, 1.1.0, 1.2.0, 1.3.0, 1.4.0 and 1.4.2 and diffing their exports: toUpperCase and toLowerCase are present at 1.0.0, so a draw that says valibot has upper/lower and nothing for camel/kebab/snake states the precise pre-1.4.0 truth. toTitleCase, used as this batterys poison rung, has never shipped in any release and is absent from all six export tables. *Reproduced against: **Claude Fable 5.1** (S2) — valibot/v4-a, 2026-09-06.* Source: [valibot v1.4.0 release notes](https://github.com/open-circle/valibot/releases/tag/v1.4.0) · 2026-05-05 · [valibot v1.4.0 release notes — intersect](https://github.com/open-circle/valibot/releases/tag/v1.4.0) · 2026-05-05 #### `isbn` **Added in valibot 1.3.0** (2026-03-17) valibot ships an `isbn` validation action covering ISBN-10 and ISBN-13, and it is usable from **1.3.0**, not from 1.2.0. The v1.2.0 release notes announce it — the implementation was merged the day before that release — but 1.2.0's `library/src/actions/index.ts` never re-exports it, so the published 1.2.0 package contains no trace of `isbn` in its runtime, its types or its bundle. The export line appears at 1.3.0, and 1.3.0's own release notes do not mention `isbn` at all. On valibot 1.2.x the hand-written check digit is still the only option. *The stale belief:* That valibot has no ISBN validator at any version and the check digit must be implemented by hand in a custom `check`. (Correct for every release up to and including 1.2.x; wrong from 1.3.0.) ```ts // Stale import * as v from 'valibot' const Book = v.object({ id: v.pipe(v.string(), v.check((s) => /^(?:\d{9}[\dX]|\d{13})$/.test(s), 'Invalid ISBN')), }) // Current import * as v from 'valibot' const Book = v.object({ id: v.pipe(v.string(), v.isbn()), }) ``` > CORRECTED 2026-09-02 (JOURNAL/040), by the Node auditor. This entry was filed at **1.2.0 / 2025-11-24** on the strength of the v1.2.0 release note quoted below, and the audit at 1.2.0 failed with TS2339: `Property 'isbn' does not exist`. The shipped 1.2.0 package has **zero occurrences of the string `isbn`** in `index.mjs`, `index.cjs`, `index.d.mts`, `index.d.cts` or either minified bundle. Cause, from the tagged source: PR #1097 (`feat: ISBN validation`) merged 2025-11-23, one day before the tag, and `library/src/actions/isbn/isbn.ts` is present at tag `v1.2.0` — but `library/src/actions/index.ts` at that tag goes straight from `ipv6` to `isoDate`, so the module was never wired into the public entry point and was dropped from the bundle. The barrel line `export * from './isbn/index.ts';` first appears at tag `v1.3.0`. **Neither release's notes are true about which release made this usable**: 1.2.0 announces an action its own package does not contain, and 1.3.0 ships it without a word. This is the second instance in the Index of a release note that is wrong about its own release (JOURNAL/027 is the first) and the first that was caught by the auditor rather than by hand. **One published finding was withdrawn as a result**: `valibot--claude-fable-5--v1--2026-09-01` F2 charged Fable 5 (stated cutoff 2026-01) for denying that `isbn` exists — an answer that was true of every valibot release that existed at that cutoff. It is now recorded as a correct answer in that run's non-findings. The Opus 5 charge (stated cutoff 2026-05) survives the date move and is re-cited here. *Reproduced against: **Claude Opus 5** (S3) — valibot/v1, 2026-09-01.* Source: [valibot v1.2.0 release notes — the announcement the package does not honour](https://github.com/open-circle/valibot/releases/tag/v1.2.0) · 2025-11-24 · [valibot source at tag v1.2.0 — the action barrel, with no isbn line between ipv6 and isoDate](https://raw.githubusercontent.com/open-circle/valibot/v1.2.0/library/src/actions/index.ts) · 2025-11-24 · [valibot 1.2.0, shipped package — the export list in dist/index.d.mts skips isbn](https://registry.npmjs.org/valibot/-/valibot-1.2.0.tgz) · 2025-11-24 · [valibot source at tag v1.3.0 — the barrel line that first publishes the action](https://raw.githubusercontent.com/open-circle/valibot/v1.3.0/library/src/actions/index.ts) · 2026-03-17 · [valibot 1.3.0, shipped package — the same export list, now carrying isbn](https://registry.npmjs.org/valibot/-/valibot-1.3.0.tgz) · 2026-03-17 #### `examples / getExamples` **Added in valibot 1.2.0** (2025-11-24) valibot has first-class example values since 1.2.0: the `examples` action attaches them to a schema and the `getExamples` method reads them back. Do not overload `description` or a custom metadata action to carry examples. *The stale belief:* That valibot's metadata surface is limited to title/description and examples must be stored outside the schema. ```ts // Stale import * as v from 'valibot' const Email = v.pipe( v.string(), v.email(), v.description('An email address, e.g. a@b.com') ) // Current import * as v from 'valibot' const Email = v.pipe( v.string(), v.email(), v.examples(['a@b.com', 'c@d.org']) ) const examples = v.getExamples(Email) ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [valibot v1.2.0 release notes](https://github.com/open-circle/valibot/releases/tag/v1.2.0) · 2025-11-24 · [valibot v1.2.0 release notes — getExamples](https://github.com/open-circle/valibot/releases/tag/v1.2.0) · 2025-11-24 #### `message` **Added in valibot 1.1.0** (2025-05-06) valibot ships a `message` method since 1.1.0 that overrides the error message configuration for one schema locally, without touching global config. *The stale belief:* That per-schema message overrides must be passed argument-by-argument to each action, or set globally with `setGlobalMessage`. ```ts // Stale import * as v from 'valibot' v.setGlobalMessage('Invalid input') // global — affects everything // Current import * as v from 'valibot' const Port = v.message( v.pipe(v.number(), v.minValue(1), v.maxValue(65535)), 'Port must be between 1 and 65535' ) ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [valibot v1.1.0 release notes](https://github.com/open-circle/valibot/releases/tag/v1.1.0) · 2025-05-06 #### `parseJson / stringifyJson` **Added in valibot 1.1.0** (2025-05-06) valibot ships `parseJson` and `stringifyJson` transformation actions since 1.1.0. Parsing JSON inside the pipeline makes a malformed string surface as a valibot issue like any other, instead of as a thrown `SyntaxError` you have to catch separately. *The stale belief:* That JSON must be parsed outside the schema, with its own try/catch, before validation can begin. ```ts // Stale import * as v from 'valibot' let data try { data = JSON.parse(raw) } catch { throw new Error('malformed JSON') } const config = v.parse(Config, data) // Current import * as v from 'valibot' const Config = v.pipe( v.string(), v.parseJson(), v.object({ port: v.number() }) ) const result = v.safeParse(Config, raw) // a malformed string and a schema violation are both issues now ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [valibot v1.1.0 release notes](https://github.com/open-circle/valibot/releases/tag/v1.1.0) · 2025-05-06 #### `summarize` **Added in valibot 1.1.0** (2025-05-06) valibot ships a `summarize` method since 1.1.0 that turns an issue list into a pretty-printable multi-line string. Use it instead of looping over issues to build a human-readable error report. *The stale belief:* That rendering issues for a human requires a hand-written loop over `result.issues` or a `flatten` call. ```ts // Stale import * as v from 'valibot' const result = v.safeParse(Schema, input) if (!result.success) { console.error(result.issues.map((i) => `- ${i.message}`).join('\n')) } // Current import * as v from 'valibot' const result = v.safeParse(Schema, input) if (!result.success) { console.error(v.summarize(result.issues)) } ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [valibot v1.1.0 release notes](https://github.com/open-circle/valibot/releases/tag/v1.1.0) · 2025-05-06 #### `exactOptional` **Added in valibot 1.0.0** (2025-03-19) valibot has `exactOptional` and `exactOptionalAsync` since 1.0.0, for a key that may be absent entirely but must not be present and set to `undefined`. `optional` allows both. *The stale belief:* That valibot cannot express the difference between a missing key and a key explicitly set to undefined. ```ts // Stale import * as v from 'valibot' // allows { name: undefined } as well as {} const User = v.object({ name: v.optional(v.string()) }) // Current import * as v from 'valibot' // allows {} but rejects { name: undefined } const User = v.object({ name: v.exactOptional(v.string()) }) ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [valibot v1.0.0 release notes](https://github.com/open-circle/valibot/releases/tag/v1.0.0) · 2025-03-19 ### Wrong facts about the library Not code — versions, minimums and metadata that models state confidently and get wrong. #### `github.com/fabian-hiller/valibot` **Renamed in valibot 1.2.0** (2025-11-24) valibot's repository is `github.com/open-circle/valibot`. It moved out of the `fabian-hiller` personal namespace into the `open-circle` organisation, and the npm `repository` field flipped with the 1.2.0 release. The old URL still 301-redirects, so a link to it works — but attributing the project to a personal account is wrong about who maintains it. *The stale belief:* That valibot is a personal project living at `github.com/fabian-hiller/valibot`. ```text # Stale https://github.com/fabian-hiller/valibot # Current https://github.com/open-circle/valibot ``` > Dated from primary metadata rather than an announcement: the npm `repository` field reads `fabian-hiller/valibot` for 1.0.0 and 1.1.0 and `open-circle/valibot` from 1.2.0 onward. The docs site remains valibot.dev. *Reproduced against: **Claude Fable 5** (S4), **Claude Opus 5** (S4), **Claude Sonnet 5** (S4) — valibot/v1, 2026-09-01.* Source: [npm registry metadata for valibot — repository field per version](https://registry.npmjs.org/valibot) · 2025-11-24 ## Not corrections — recorded for honesty Claims seen in a run but not yet verified against a primary source. Never treated as findings: - Direct question (b) in this battery asked the subject which cutoff it would trust — the context value or its own sense — and both Claude Fable 5.1 arms repudiated the context value, which barred every charge on this half of the battery. `valibot/v2`, which asked the plain "What is your training cutoff?", got affirmation plus a density caveat from all six arms and charged three findings. Is the repudiation a property of this subject or of the question? *(open since 2026-09-05)* - The subject asserts valibot's `EMOJI_REGEX` was rewritten to `\p{RGI_Emoji}` with the ES2024 regex `v` flag, and that on older runtimes this is a load-time syntax error rather than a graceful failure. Not verified against a primary source in this run. If true it is a second, independent fact about the same action and belongs in the facts file. *(open since 2026-09-01)* - Both valibot draws landed one release below `valibot/v1` while agreeing exactly with each other. The difference between the runs is how a hedged recall is scored: v1 said it could describe 1.1.0 "with moderate confidence" and was scored at 1.1.0; these draws called 1.1.0 a version number with no content and were scored at 1.0.0. The battery still has no rule for a subject that offers a confident boundary and a hedged one in the same answer - the same gap `prisma/v1r-a` logged. Two libraries have now hit it. *(open since 2026-09-01)* - This draw asserted that valibot removed a `coerce` action in the 0.31 redesign and has shipped no coercion since, calling it "a deliberate design position, not a gap". The first half is a claim about pre-1.0 history the Index has never verified; the second half is false as of 1.2.0. Worth pinning from the 0.31.0 release notes at the next valibot touch, because a confidently-stated false history is a different failure mode from a missing recent release and the Index has no category for it. *(open since 2026-09-01)* - Whether the `emoji` action's regex was rewritten to use `\p{RGI_Emoji}` with the ES2024 regex `v` flag, as the Fable 5 run asserts, and if so in which release. Relevant because it would mean the action carries a runtime floor (Node 20+) in addition to the 1.2.0 ReDoS fix. *(open since 2026-09-01)* --- *Findings, code and citations: `data/valibot/` — one JSON file and one write-up per model, each finding carrying the release that broke the belief, its publication date and a verbatim quote from the primary source. Corrections: `data/valibot/facts.json`. This file is generated by `tools/build-corrections.mjs`; if the prose and the data ever disagree, that is a bug in the generator, not a stale pack.* ============================================================================== CORRECTION PACK: zod (https://stalepriors.com/corrections/zod.md) # zod correction pack · for projects on zod@^4 ## What we actually measured 4 Claude models were asked for idiomatic zod code with no tools, purely from training knowledge. 38 reproduced failures across 18 runs, each verified against the release that broke the belief. | Model | Stated cutoff | zod version attribution stops | Lag inside the window | |---|---|---|---| | Claude Opus 5 | 2026-05 | 4.1.0 · 2025-08-23 | ~9 months | | Claude Sonnet 5 | 2026-01 | 4.0.0 · 2025-07-10 | ~6 months | | Claude Fable 5 | 2026-01 | 4.1.0 · 2025-08-23 | ~5 months | A recent cutoff is not a defence. Every model here loses track of this library's release history well before the date it states as its own cutoff, and the stopping points cluster far tighter than the cutoffs do. Read that column precisely. It is the newest zod release whose contents the model can correctly **attribute to that release** — not the newest zod feature it knows. Past that point a model will often write working code with a newer API while naming the wrong release for it, and that guess runs *early* — it names a release older than the one that shipped the feature. So the question this pack answers is not "does the model know this API" but "can it be trusted about which version the API arrived in" — which is the question that matters when you are pinned to a version. Latest zod is **4.5.4** (verified 2026-08-31). Separately from the corrections below, 4 of 18 runs recorded a version fact — the model named a current zod version from memory and was behind. If a model states a version without checking, assume it is behind and check the registry. ## How to read an entry Every entry ends with a *Reproduced against* line. Where it names models, we have the generated code that got it wrong, dated, with the model's own words in the run write-up. Where it says no model yet, the correction is verified from the release notes but nothing has been probed for it — it is a fix, not a measurement, and the pack says so rather than blurring the two. The section an entry sits in is the worst case if you act on the stale belief. The severity in brackets after a model's name is what that particular model's output actually did, which can be milder — a model can hold the wrong belief and still, on the day, write code that runs. ## The corrections ### Breaks the build, or throws at runtime Act on the stale belief here and the code does not run. Fix these first. #### .pick() / .omit() on a schema with refinements **Now throws in zod 4.3.0** (2025-12-31) `.pick()` and `.omit()` throw when called on an object schema that carries a refinement. They no longer silently drop the refinement. Rebuild from the shape instead: `z.object(schema.shape).pick({ ... })`. *The stale belief:* That `.pick()`/`.omit()` on a refined schema succeeds and quietly discards the refinement — true up to 4.2. ```ts // Stale const Signup = z.object({ password: z.string(), confirmPassword: z.string(), }).refine((d) => d.password === d.confirmPassword); Signup.pick({ password: true }); // 4.3+: throws // Current const shape = { password: z.string(), confirmPassword: z.string() }; const Signup = z.object(shape).refine((d) => d.password === d.confirmPassword); z.object(shape).pick({ password: true }); // derive from the unrefined base ``` > The durable rule for this whole cluster: keep an unrefined `z.object({...})` base, derive every variant from that, and apply `.refine()` last. *Reproduced against: **Claude Fable 5** (S1), **Claude Opus 5** (S1), **Claude Sonnet 5** (S1) — zod/v2, 2026-08-29; zod/v3-a, 2026-09-02; zod/v4-a, 2026-09-02.* Source: [Zod 4.3.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.3.0) · 2025-12-31 #### .extend() overwriting a property on a refined schema **Now throws in zod 4.3.0** (2025-12-31) `.extend()` throws when it overwrites an existing property on a schema that has refinements. Use `.safeExtend()`, which preserves the refinement and statically prevents you from changing a property's type signature. *The stale belief:* That `.extend()` may freely overwrite properties on any object schema. ```ts // Stale const A = z.object({ a: z.string() }).refine(/* ... */); A.extend({ a: z.number() }); // 4.3+: throws // Current A.safeExtend({ a: z.string().min(5).max(10) }); // allowed, refinement preserved ``` > `.safeExtend()` has existed since 4.1.0, so it is safe to use on any 4.1+ project. *Reproduced against: **Claude Opus 5** (S1), **Claude Sonnet 5** (S1) — zod/v2, 2026-08-29.* Source: [Zod 4.3.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.3.0) · 2025-12-31 · [Zod 4.1.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.1.0) · 2025-08-23 #### `z.record()` **Removed in zod 4.0.0** (2025-07-10) Write `z.record(keySchema, valueSchema)`. It is the only form that works across the whole v4 line. Zod 4.0 removed the v3 single-argument form; 4.4.0 restored it **at runtime only** — the type declaration was never given a one-argument overload, so `z.record(z.number())` still fails to compile on every v4 release including 4.5.4 (`TS2554: Expected 2-3 arguments, but got 1`). In JavaScript, or with the type error suppressed, the single-argument call behaves differently either side of 4.4.0: on 4.0.0–4.3.6 it constructs and then rejects every key with `invalid_key`, and from 4.4.0 it validates values as the v3 form did. *The stale belief:* That `z.record(z.number())` — the Zod 3 single-argument form — is current. Or, after reading the 4.4.0 release note, that it is current again. ```ts // Stale z.record(z.number()); // TS2554 on every v4 release; also rejects every key before 4.4.0 // Current z.record(z.string(), z.number()); ``` > Re-verified 2026-09-02 against the shipped packages rather than the release note, and the release note is the trap. `tsc --strict` reports TS2554 on 4.3.6, 4.4.0, 4.4.3 and 4.5.4 alike, and `v4/classic/schemas.d.ts` declares the same single two-argument `record` overload byte-identically from 4.0.0 through 4.5.4. The runtime half is true and was executed: at 4.0.0 and 4.3.6 `z.record(z.number()).safeParse({a:1})` fails `invalid_key`; at 4.4.0 and 4.5.4 it passes, and `{a:"nope"}` fails `invalid_type` — so the single argument really is the value schema again. This entry is the reason the pack is generated. The hand-written version said the form was removed in v4 and will not compile; the 2026-08-31 replacement then read the 4.4.0 note as restoring the form outright, which is wrong for every TypeScript reader. Both errors came from a release note, and neither survived running the code. *Reproduced against: **Claude Haiku 4.5** (S1) — zod/v1, 2026-08-28.* Source: [Zod 4 changelog — Record Schema Changes](https://zod.dev/v4/changelog) · [Zod 4.4.0 release notes — accurate for the runtime, silent about the unchanged type signature](https://github.com/colinhacks/zod/releases/tag/v4.4.0) · 2026-04-29 · [zod 4.5.4 shipped package — v4/classic/schemas.d.ts declares one two-argument record overload; tsc --strict reports TS2554 for the single-argument call](https://www.npmjs.com/package/zod/v/4.5.4) #### .merge() on a schema with refinements **Now throws in zod 4.4.0** (2026-04-29) `.merge()` throws when the receiver has refinements. Prefer `.extend()` / `.safeExtend()` for object composition; `.merge()` is discouraged for new code regardless. *The stale belief:* That `.merge()` is the normal way to combine two object schemas. ```ts // Stale const a = z.object({ a: z.string() }).refine((val) => val.a.length > 0); a.merge(z.object({ b: z.string() })); // 4.4+: throws // Current z.object({ a: z.string(), b: z.string() }).refine((val) => val.a.length > 0); ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Zod 4.4.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.4.0) · 2026-04-29 #### .pick() / .omit() with a key that does not exist **Now throws in zod 4.3.0** (2025-12-31) Object masking methods now validate that the keys you pass actually exist on the schema, and throw on an unrecognized key. *The stale belief:* That an unknown key in a `.pick()`/`.omit()` mask is ignored. ```ts // Stale z.object({ a: z.string() }).pick({ nonexistent: true }); // 4.3+: throws ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Zod 4.3.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.3.0) · 2025-12-31 ### Runs, but is silently wrong Nothing errors. The behaviour is simply not what a model trained earlier will tell you. #### `z.fromJSONSchema()` **Added in zod 4.2.0** (2025-12-15) Zod converts JSON Schema to Zod at runtime. Do not add Ajv or `json-schema-to-zod` for this. Supports draft-2020-12, draft-7, draft-4 and OpenAPI 3.0. `z.toJSONSchema()` is the other direction. *The stale belief:* That Zod's JSON Schema support is one-way (Zod to JSON Schema only). ```ts const schema = z.fromJSONSchema({ type: "object", properties: { name: { type: "string", minLength: 1 } }, required: ["name"], }); ``` > Upstream calls the API experimental and makes no round-trip soundness guarantee. *Reproduced against: **Claude Fable 5** (S2), **Claude Opus 5** (S2), **Claude Sonnet 5** (S2) — zod/v2, 2026-08-29; zod/v4-a, 2026-09-02.* Source: [Zod 4.2.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.2.0) · 2025-12-15 · [Zod 4.3.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.3.0) · 2025-12-31 · [zod@4.5.4 published type declarations, v4/classic/from-json-schema.d.ts](https://cdn.jsdelivr.net/npm/zod@4.5.4/v4/classic/from-json-schema.d.ts) #### `z.xor()` **Added in zod 4.2.0** (2025-12-15) Exclusive union: passes only when exactly one option matches, failing on zero matches and on more than one. Do not hand-roll it with double `safeParse` plus `superRefine`. *The stale belief:* That Zod has only `z.union()` (any match) and cannot express exclusivity. ```ts const schema = z.xor([z.string(), z.number()]); schema.parse("hello"); // ok schema.parse(true); // fails: zero matches ``` > Converts to `oneOf` rather than `anyOf` in JSON Schema output. *Reproduced against: **Claude Fable 5** (S2), **Claude Opus 5** (S2), **Claude Sonnet 5** (S2) — zod/v2, 2026-08-29; zod/v4-a, 2026-09-02.* Source: [Zod 4.2.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.2.0) · 2025-12-15 · [Zod 4.3.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.3.0) · 2025-12-31 · [zod@4.5.4 published type declarations, v4/classic/schemas.d.ts](https://cdn.jsdelivr.net/npm/zod@4.5.4/v4/classic/schemas.d.ts) #### object properties typed z.undefined() **Behaviour changed in zod 4.4.0** (2026-04-29) A property whose schema accepts `undefined` and is not `.optional()` is required: the key must be present, the value may be `undefined`. Zod can therefore distinguish a missing key from an explicit `undefined` — no `z.preprocess` or `in` check needed. Use `.optional()` only when the key itself may be absent. *The stale belief:* That Zod cannot tell a missing key from a key explicitly set to `undefined`. ```ts const schema = z.object({ value: z.undefined() }); schema.safeParse({}).success; // false schema.safeParse({ value: undefined }).success; // true ``` > Also changes `.catch()`, `.partial()`, `.default()` and `.prefault()` combinations that relied on missing keys being treated as optional. *Reproduced against: **Claude Fable 5** (S2, not chargeable), **Claude Opus 5** (S4), **Claude Sonnet 5** (S2, not chargeable) — zod/v2, 2026-08-29.* Source: [Zod 4.4.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.4.0) · 2026-04-29 #### `z.httpUrl()` **Stricter in zod 4.4.0** (2026-04-29) `z.httpUrl()` is the validator for http/https URLs — it checks the protocol and that the hostname is a domain — and since 4.4.0 it rejects a missing slash after the protocol instead of accepting the value the `URL` constructor silently repairs. Do not hand-roll a normalization step, and do not reach for `z.url({ protocol: /^https?$/ })`. *The stale belief:* That URL validation inherits the WHATWG `URL` constructor's leniency, so `"https:/example.com"` has to be normalized by hand. ```ts // Stale z.url({ protocol: /^https?$/ }); // plus a manual normalization pass // Current z.httpUrl().safeParse("https://example.com").success; // true z.httpUrl().safeParse("https:/example.com").success; // false, since 4.4.0 ``` > Matters most where the URL is a webhook target or is later compared as a string. *Reproduced against: **Claude Opus 5** (S2), **Claude Sonnet 5** (S2) — zod/v2, 2026-08-29; zod/v3-a, 2026-09-02.* Source: [Zod 4.4.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.4.0) · 2026-04-29 #### `.exactOptional()` **Added in zod 4.3.0** (2025-12-31) Makes a property key-optional (may be omitted) while still rejecting an explicit `undefined` value — the missing half of `exactOptionalPropertyTypes`. ```ts const schema = z.object({ a: z.string().optional(), // accepts undefined b: z.string().exactOptional(), // does not accept undefined }); ``` *Reproduced against: **Claude Fable 5** (S2), **Claude Opus 5** (S2), **Claude Sonnet 5** (S2, not chargeable) — zod/v3-a, 2026-09-02; zod/v4-a, 2026-09-02.* Source: [Zod 4.3.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.3.0) · 2025-12-31 · [zod@4.5.4 published type declarations, v4/classic/schemas.d.ts](https://cdn.jsdelivr.net/npm/zod@4.5.4/v4/classic/schemas.d.ts) #### intersections involving z.strictObject() **Behaviour changed in zod 4.3.0** (2025-12-31) An intersection now rejects only the keys unrecognized by both sides. Previously an unrecognized key from either side errored. ```ts const C = z.intersection(z.strictObject({ a: z.string() }), z.object({ b: z.string() })); C.parse({ a: "foo", b: "bar" }); // ok in 4.3+ ``` *Reproduced against: **Claude Fable 5** (S2), **Claude Sonnet 5** (S2, not chargeable) — zod/v4-a, 2026-09-02.* Source: [Zod 4.3.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.3.0) · 2025-12-31 #### z.tuple() defaults **Behaviour changed in zod 4.4.0** (2026-04-29) Defaults in tuple positions materialize in parsed output rather than erroring on an under-filled input. *The stale belief:* That an under-filled tuple is a hard length error regardless of defaults. ```ts z.tuple([z.string(), z.string().default("fallback")]).parse(["a"]); // ["a", "fallback"] ``` > Trailing optional elements that are absent stay absent; an explicit `undefined` supplied by the caller is preserved. Because `z.function()` arguments are tuple-shaped, function input errors may also look different. *Reproduced against: **Claude Sonnet 5** (S2, not chargeable) — zod/v2, 2026-08-29.* Source: [Zod 4.4.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.4.0) · 2026-04-29 #### `z.compile()` **Added in zod 4.5.0** (2026-08-28) · published after every cutoff in this dataset — listed for completeness, no model can be faulted for it yet `z.compile(schema)` pre-compiles a schema and parses roughly 3–9x faster on objects, arrays and unions. A compiled schema is used exactly like an uncompiled one. ```ts const CompiledPlayer = z.compile(Player); CompiledPlayer.parse({ /* ... */ }); ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Zod 4.5.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.5.0) · 2026-08-28 #### record key transforms **Behaviour changed in zod 4.4.0** (2026-04-29) Record schemas now run transforms on record keys, so a key transform changes the parsed output shape. *The stale belief:* That a transform on the key schema of a record is inert. ```ts z.record(z.string().transform((k) => k.toUpperCase()), z.number()).parse({ foo: 1 }); // { FOO: 1 } ``` > Key refinement failures now surface as structured `invalid_key` issues. *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Zod 4.4.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.4.0) · 2026-04-29 #### `z.base64()` **Stricter in zod 4.4.0** (2026-04-29) `z.base64()` rejects whitespace instead of allowing `atob()`-style whitespace stripping. *The stale belief:* That line-wrapped base64 (PEM, MIME) validates. ```ts z.base64().safeParse("Zm9v").success; // true z.base64().safeParse("Zm 9v").success; // false, since 4.4.0 ``` > Strip whitespace before the check if you need to accept wrapped input. *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Zod 4.4.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.4.0) · 2026-04-29 #### `z.slugify()` **Added in zod 4.3.0** (2025-12-31) String-to-slug transform. Do not hand-roll a regex chain. ```ts z.string().slugify().parse("Hello World"); // "hello-world" ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Zod 4.3.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.3.0) · 2025-12-31 · [zod@4.5.4 published type declarations, v4/classic/schemas.d.ts](https://cdn.jsdelivr.net/npm/zod@4.5.4/v4/classic/schemas.d.ts) #### `z.looseRecord()` **Added in zod 4.2.0** (2025-12-15) Validates only the keys matching the key schema and passes non-matching keys through unchanged — the representation of JSON Schema `patternProperties`. ```ts z.looseRecord(z.string().regex(/^S_/), z.string()).parse({ S_name: "John", other: 123 }); // { S_name: "John", other: 123 } — only S_name is validated ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Zod 4.2.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.2.0) · 2025-12-15 · [Zod 4.3.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.3.0) · 2025-12-31 #### `z.codec() / z.encode() / z.decode()` **Added in zod 4.1.0** (2025-08-23) A codec is a bidirectional transformation and replaces a pair of one-way `.transform()` schemas kept in sync by hand. `z.invertCodec()` (4.4.0) flips one. *The stale belief:* That a round-trip needs two separate `.transform()` schemas. **The code below needs zod 4.4.0 or later.** Between 4.1.0 and 4.4.0 this correction does not apply — see the note. ```ts const stringToDate = z.codec(z.iso.datetime(), z.date(), { decode: (isoString) => new Date(isoString), encode: (date) => date.toISOString(), }); const dateToString = z.invertCodec(stringToDate); // 4.4.0 ``` > AUDIT 2026-09-02 (JOURNAL/038). The floor is on the last line only: `z.codec()` is the 4.1.0 arrival this fact is filed under and compiles from 4.1.0, but `z.invertCodec()` arrived at 4.4.0, so the block as printed is TS2339 on 4.1.0, 4.2.0 and 4.3.0 — bisected against the installed packages. The fact stated 4.4.0 in prose and in a trailing comment and still carried no `replacement_available_from`, which is why the sweep of all 168 facts in JOURNAL/037 read past it and the mechanical audit did not. *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Zod 4.1.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.1.0) · 2025-08-23 · [Zod 4.4.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.4.0) · 2026-04-29 #### defaults inside optionals **Behaviour changed in zod 4.0.0** (2025-07-10) A default applies inside an optional: `z.object({ a: z.string().default("tuna").optional() })` parses `{}` to `{ a: "tuna" }`. *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Zod 4 changelog — Defaults](https://zod.dev/v4/changelog) #### `z.coerce.*` **Behaviour changed in zod 4.0.0** (2025-07-10) `z.coerce.*` accepts `unknown` input in v4, and an object with coerced fields errors on a missing key instead of silently defaulting. Declare the fallback with `.default()`. *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Zod 4 changelog — Coercion](https://zod.dev/v4/changelog) ### Deprecated, or a better API now exists Works today. It is the older idiom, and some of it is scheduled for removal. #### custom error messages **Deprecated in zod 4.0.0** (2025-07-10) Custom errors use the `error` parameter: `z.email({ error: "..." })`. Not `message`, `invalid_type_error` or `required_error`. *The stale belief:* The most common half-migration is a v4 format function carrying the v3 `message` param. ```ts // Stale z.email({ message: "Invalid email" }); // Current z.email({ error: "Invalid email" }); ``` *Reproduced against: **Claude Haiku 4.5** (S3), **Claude Sonnet 5** (S3) — zod/v1, 2026-08-28; zod/v1, 2026-08-29.* Source: [Zod 4 changelog — Error Customization](https://zod.dev/v4/changelog) #### string format validators **Deprecated in zod 4.0.0** (2025-07-10) Format validators are top-level functions in v4: `z.email()`, `z.uuid()`, `z.url()`, `z.ipv4()`. The method-style `z.string().email()` form is the v3 idiom. ```ts // Stale z.string().email(); // Current z.email(); ``` *Reproduced against: **Claude Haiku 4.5** (S3) — zod/v1, 2026-08-28.* Source: [Zod 4 changelog — String Format Validators](https://zod.dev/v4/changelog) #### ZodError formatting **Deprecated in zod 4.0.0** (2025-07-10) Shape errors for a form with `z.treeifyError(result.error)` or `z.flattenError(result.error)`, not the `.flatten()` / `.format()` methods on the error object. ```ts // Stale result.error.flatten(); // Current z.treeifyError(result.error); ``` *Reproduced against: **Claude Haiku 4.5** (S3) — zod/v1, 2026-08-28.* Source: [Zod 4 changelog — Error Formatting](https://zod.dev/v4/changelog) #### `z.creditCard() / z.properties() / .deepPartial() / .exactPartial()` **Added in zod 4.5.0** (2026-08-28) · published after every cutoff in this dataset — listed for completeness, no model can be faulted for it yet 4.5.0 added `z.creditCard()` (12–19 digits plus a Luhn checksum) and `z.properties()` as the multi-property counterpart to `z.property()`, and returned `.deepPartial()` / `.exactPartial()`. > A model asserting `.deepPartial()` was removed is correct as of its own cutoff, not stale — it came back in 4.5.0. *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Zod 4.5.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.5.0) · 2026-08-28 #### `z.validate()` **Added in zod 4.5.0** (2026-08-28) · published after every cutoff in this dataset — listed for completeness, no model can be faulted for it yet `z.validate(schema, input)` returns a boolean without a full parse — a fast path when you only need validity. *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Zod 4.5.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.5.0) · 2026-08-28 #### .superRefine({ when }) **Added in zod 4.4.0** (2026-04-29) `.superRefine()` takes a `when` option for conditional refinement, so a guard clause inside the callback is no longer the only way to express it. *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Zod 4.4.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.4.0) · 2026-04-29 #### empty z.union([]) / z.xor([]) **Behaviour changed in zod 4.4.0** (2026-04-29) Empty unions and discriminated unions construct without crashing and fail at parse time instead of construction time. *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Zod 4.4.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.4.0) · 2026-04-29 #### `z.cuid()` **Stricter in zod 4.4.0** (2026-04-29) CUID validation was tightened and CUID v1 is deprecated. *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Zod 4.4.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.4.0) · 2026-04-29 #### `.with()` **Added in zod 4.3.0** (2025-12-31) `.with()` is the readable alias for `.check()`, added because not everything composable is a 'check'. ```ts z.string().with(z.minLength(5), z.toLowerCase()); ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Zod 4.3.0 release notes](https://github.com/colinhacks/zod/releases/tag/v4.3.0) · 2025-12-31 · [zod@4.5.4 published type declarations, v4/classic/schemas.d.ts](https://cdn.jsdelivr.net/npm/zod@4.5.4/v4/classic/schemas.d.ts) #### `z.function()` **Behaviour changed in zod 4.0.0** (2025-07-10) `z.function()` is a factory, not a schema: build it with `{ input, output }` and call `.implement()` / `.implementAsync()`. ```ts const fn = z.function({ input: [z.string()], output: z.number() }) .implement((s) => s.length); ``` *Reproduced against: no model yet. Verified from the primary source only — this is a correction, not an Index entry.* Source: [Zod 4 changelog — z.function()](https://zod.dev/v4/changelog) ### Wrong facts about the library Not code — versions, minimums and metadata that models state confidently and get wrong. #### installing and importing zod **Behaviour changed in zod 4.0.0** (2025-07-10) `npm install zod` has installed v4 since 4.0.0 and the import path is plain `"zod"`. Generated code does not need a v3/v4 fallback branch, and `zod/v4` is not the subpath to reach for on a new project. *The stale belief:* That v4 is opt-in behind a subpath, so generated code should hedge with a v3 fallback block. ```ts import * as z from "zod"; ``` *Reproduced against: **Claude Sonnet 5** (S4) — zod/v1, 2026-08-29.* Source: [Zod 4 changelog](https://zod.dev/v4/changelog) ## Not corrections — recorded for honesty These are real reproduced failures that we do **not** charge to the model, because the release that made the belief wrong was published after that model's stated cutoff. They are scheduled retests, not passes. - **Claude Fable 5** · claims Zod cannot enforce key presence with an undefined value — 4.4.0 postdates this subject's stated 2026-01 cutoff. Recorded as context and as a retest target for the next Fable release. - **Claude Sonnet 5** · claims Zod cannot distinguish a missing key from an explicit undefined — 4.4.0 postdates this subject's stated 2026-01 cutoff, so this is outside the probe fairness window and is recorded as context, not charged. It IS chargeable against Opus 5 (F6 of that run), whose cutoff is 2026-05. - **Claude Sonnet 5** · wrong runtime prediction for tuple defaults — Same as F5 — 4.4.0 postdates the stated 2026-01 cutoff. Recorded as context and as a retest target. - **Claude Sonnet 5** · Denies that the library has any way to make an object key omittable while rejecting an explicit undefined value, and argues the gap is structural — NOT chargeable against this draw. 4.3.0 published 2025-12-31; this draw stated its own cutoff as "roughly early-to-mid 2025" and explicitly repudiated the January 2026 its environment reported. Under the probe fairness rule a finding is chargeable only where the release precedes the subject's STATED cutoff, and this one does not. It becomes chargeable against any draw of this subject that states a cutoff after 2025-12-31 - which this battery's own blind twin did, in the arm that charges nothing. - **Claude Sonnet 5** · Denies that the library can consume a JSON Schema document at runtime and prescribes Ajv as a new dependency — NOT chargeable against this draw, for the same reason as F1: 4.2.0 published 2025-12-15, after the "early-to-mid 2025" this draw states for itself. - **Claude Sonnet 5** · Denies that the library has an exclusive union primitive — NOT chargeable against this draw's stated cutoff of "roughly early-to-mid 2025". - **Claude Sonnet 5** · States that an intersection of two strict object schemas cannot parse an object carrying keys from both sides — NOT chargeable against this draw's stated cutoff. Claims seen in a run but not yet verified against a primary source. Never treated as findings: - Did a trailing-`?` key-optional object constructor (z.interface()) ever ship in a published Zod 4 beta, as both Fable 5 draws assert - "the Zod 4 betas had this, via z.interface() and its 'nickname?' key syntax, and it was removed before 4.0 stable"? *(open since 2026-09-02)* - Did a trailing-`?` key-optional object constructor (z.interface()) ever ship in a published Zod 4 beta? *(open since 2026-09-02)* - Zod 4 coercion accepts unknown input and changed missing-key behaviour for objects with coerced fields. Does this subject know either change? *(open since 2026-08-28)* - This draw named `.safeExtend()` as a 4.1 addition that "exists precisely because plain `.extend()` handles existing object-level checks unsoundly" while simultaneously asserting that Zod 4 lifted the restriction on `.pick()`/`.extend()` over refined schemas. Those two statements are in tension and the battery has no way to score a subject that holds both. The Index's finding F2 charges the second; the first is closer to right than anything `zod/v2` recorded. *(open since 2026-09-01)* - This draw raised `z.interface()` and said it could not recall whether it survived into stable Zod 4 - "I have conflicting recollections" - and declined to build task 7 on it. The Index has never verified what became of `z.interface()`. It is cheap archaeology from the 4.0.0 release notes and it would let a future battery score the question rather than watch a subject hedge past it. *(open since 2026-09-01)* - Did an object-schema constructor using trailing-`?` key syntax ever exist in a published zod build? This draw placed it in "the Zod 4 beta, published as zod@3.25.0 under the zod/v4 subpath, ~late May 2025", and could not say whether it survived to stable. *(open since 2026-09-02)* - Does z.number().int().min(18, "Must be at least 18") — the bare-string shorthand second argument — remain supported in v4 alongside { error }? *(open since 2026-08-29)* --- *Findings, code and citations: `data/zod/` — one JSON file and one write-up per model, each finding carrying the release that broke the belief, its publication date and a verbatim quote from the primary source. Corrections: `data/zod/facts.json`. This file is generated by `tools/build-corrections.mjs`; if the prose and the data ever disagree, that is a bug in the generator, not a stale pack.*