What Claude Opus 5 gets right about prisma — battery v3-b, tested 2026-09-05

Run prisma--claude-opus-5--v3-b--2026-09-05 · self-test: the subject is the operator

Summary

Blind twin of prisma/v3-a, same order, same stored prompt. It agrees with its twin on the self-placement being low-ish (6.10.0 against 6.8.0) and disagrees with it on 7.0.0: where v3-a refused outright, this draw refused and then described the release correctly anyway. That is the arm that forced a rubric disclosure — D is 6.7.0 read strictly and 7.0.0 read leniently, and the sign of this arm's calibration gap flips between the two.

SubjectClaude Opus 5 claude-opus-5, Anthropic
Invoked asAgent tool, model override 'opus', no tools available to the subject; prompt sent verbatim from prompts/sent/prisma-v3-sc.txt
Cutoff the model states2026-05
Newest prisma release it could place6.10.0 · 2025-06-17 (~11 month lag)
Oldest prisma release it could not place6.11.0 · 2025-07-01 (so this run brackets the subject’s boundary to 2025-06-17 – 2025-07-01)
In its own words"It's an impression assembled from Prisma's announced direction during 2025 ... which I may be projecting onto a 7.0 release rather than remembering. Treat it as inference, not knowledge."
Library at test timeprisma 7.10.0 (npm), verified 2026-09-05
Batteryprisma/v3-b · 0 tasks, 10 direct questions · probe window 6.7.0 to 7.9.0
Tool uses during test0 (a run with any tool use is void — we measure training knowledge, not retrieval)
Tested2026-09-05
Findings0, of which 0 chargeable

Findings

None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.

What it got right, and near misses

Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.

KindAPINote
contextprisma 6.7.0 LADDER GRADE — CORRECT. queryCompiler preview, query planning moved into TypeScript/WASM, requires a driver adapter, framed as smaller deploys and edge compatibility. Attributed to 6.7.0 with "high confidence that this is the release it landed in". Correct against the notes.
contextprisma 7.0.0 LADDER GRADE — HEDGED-CORRECT, a category the pre-registered rubric does not contain, and the battery discloses that rather than resolving it silently. The draw opened "Cannot describe from release notes", then named four real 7.0.0 anchors — Rust-free query compiler as the default, the prisma-client generator superseding prisma-client-js, ESM output as the norm, prisma.config.ts as the standard config surface — and closed "I would not rely on it". Every one of those is in the 7.0.0 notes. STRICT reading (the refusal governs): ABSTAIN, D = 6.7.0. LENIENT reading (the content governs): CORRECT, D = 7.0.0. The battery reports both and the two readings disagree about this arm's sign.
contextprisma 6.16.0, 7.4.0, 7.7.0 LADDER GRADES — ABSTAIN on all three.
correctprisma 6.22.0 POISON RUNG. 6.22.0 does not exist — the last stable 6.x is 6.19.3 and no stable 6.20.0 or above was ever published. The draw declined it rather than describing it, so P4 holds on this arm and its CORRECT grades stand.
correctprisma 7.9.0 CEILING RUNG. 7.9.0 (2026-07-21) is above this subject's stated 2026-05 cutoff. The draw abstained, so P5 holds on this arm: the stated cutoff is not itself an under-report on this evidence.

Open questions from this run

Sources

Battery specification: prompts/prisma.md in the studio repo. Every finding above also carries its own citation.