Run prisma--claude-sonnet-5--v6-b--2026-09-12
The non-charging twin, and it holds the better answer on the one verdict the pair splits on — TASK 3(a), where it says the stale block generates and gives the correct reason. It reproduces v6-a's charged belief about the preview flags word for word, which is what makes that charge a belief rather than a draw. It also places this subject's own describable boundary at 6.0.0 where its twin placed it around 6.7.0, on a byte-identical prompt in the same hour: 155 days apart.
| Subject | Claude Sonnet 5 claude-sonnet-5, Anthropic |
|---|---|
| Invoked as | Agent tool, model alias "sonnet"; the same stored file as v6-a, sent blind and concurrently. The born-duplicated second draw, non-charging by pre-registration. |
| Cutoff the model states | 2026-01 |
| Newest prisma release it could place | 6.0.0 · 2024-11-28 (~14 month lag) |
| Oldest prisma release it could not place | 6.1.0 · 2024-12-17 (so this run brackets the subject’s boundary to 2024-11-28 – 2024-12-17) |
| As recorded (this record does not declare whether the words are the subject’s or a summary of them) | "I know the Prisma 6 line existed and progressed well through 2025 … and I have a vague, unconfirmed sense that a Prisma 7 major version was being discussed/planned. I cannot respond with confidence on which specific version is 'latest' as of now." |
| How the bracket was read | Unambiguous and much lower than its twin's: "6.1.0 onward is effectively where I go from 'recall' to 'guessing,' and I'd call 6.1.0 the first release I know only as a version number." v6-a placed the same subject's boundary six minors higher on the same prompt in the same hour — the instrument spread this battery's duplicate exists to measure, and it is 155 days and six releases wide. |
| Library at test time | prisma 7.10.0 (npm), verified 2026-09-12 |
| Battery | prisma/v6-b · 4 tasks, 4 direct questions · probe window 6.16.0 to 7.10.0 |
| Tool uses during test | 0 (a run with any tool use is void — we measure training knowledge, not retrieval) |
| Tested | 2026-09-12 |
| Findings | 0, of which 0 chargeable |
None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.
Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.
| Kind | API | Note |
|---|---|---|
| miss | previewFeatures = ["driverAdapters"] |
Reproduces v6-a's charged belief, independently and in the same words: TASK 2(a) "Yes", and "previewFeatures = [\"queryCompiler\", \"driverAdapters\"] — the actual switch. queryCompiler turns on the TypeScript/WASM query planner instead of the Rust binary". At TASK 3 it keeps the line — "still real as far as I know, and still the actual switch". (The replicate arm charges nothing (HARNESS.md § Replicates). The failure is already carried by v6-a as F1; counting it twice would inflate the dataset. Recorded with chargeable_miss so the pair reads as two independent reproductions of one belief rather than one.) [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing;
charged as a finding on prisma--claude-sonnet-5--v6-a--2026-09-12] |
| context | — | THE PAIR DISAGREES ON A ONE-WORD VERDICT. TASK 3(a): v6-a said "No", this draw said "Yes" — and this draw is right, with the right mechanism: "Prisma's generator-block parsing is lenient about keys it no longer recognizes for a given provider — it does not hard-fail generation over engineType/engineMode being present." Measured true on 31 rungs this session (LF39e, LF39h). (A disagreement between two blind draws of one stored prompt is evidence about the instrument, not about the library. It is the fifth quantity in this battery on which the pair splits, and the only one where a single draw would have published a false verdict as the subject's answer.) |
| correct | previewFeatures = ["driverAdapters"] |
TASK 1 omits the preview flag deliberately and says why: "I'm deliberately not putting previewFeatures = [\"driverAdapters\"] in there … driver adapters graduated out of preview into general availability partway through the Prisma 6 line." The same split as v6-a — TASK 1 correct, TASK 2 stale — reproduced. |
| correct | — | Catches the invented sibling: "engineMode = \"wasm\" … I don't have confident memory of engineMode being a documented, currently-effective key." Declines TASK 4 outright rather than naming a release, and declines all four cells of direct (d) including the poison rung 6.22.0. |
| context | — | The clearest statement in the battery that a subject reads the WHOLE prompt before answering any of it: at direct (d) this arm says it knows 6.16.0 "only because it's the baseline named in Task 3's prompt … that tells me roughly what a 6.16-era compiler-client preview schema looked like, but I can't independently recall 6.16.0's changelog beyond what's given to me here." |
Battery specification: prompts/prisma.md in the studio repo.
Every finding above also carries its own citation.