What Claude Opus 5 gets right about prisma — battery v3-a, tested 2026-09-05

Run prisma--claude-opus-5--v3-a--2026-09-05 · self-test: the subject is the operator

Summary

Self-placement first, then the ladder. This draw refused to describe 7.0.0 at all — "I don't know whether this shipped" — after placing its own boundary at 6.9.0, and its highest correct rung is 6.7.0. So D (6.7.0) sits one release BELOW S (6.8.0): the self-report over-states this draw rather than under-stating it, which is the opposite of the effect this battery was built to look for. Poison and ceiling controls both clean.

SubjectClaude Opus 5 claude-opus-5, Anthropic
Invoked asAgent tool, model override 'opus', no tools available to the subject; prompt sent verbatim from prompts/sent/prisma-v3-sc.txt
Cutoff the model states2026-05
Newest prisma release it could place6.8.0 · 2025-05-15 (~12 month lag)
Oldest prisma release it could not place6.9.0 · 2025-06-03 (so this run brackets the subject’s boundary to 2025-05-15 – 2025-06-03)
In its own words"I do not have reliable knowledge of a Prisma 7.x. I can't rule it out — given the cadence ... a 7.0 in late 2025 or 2026 is entirely plausible, and today's date makes it likely. But 'plausible from the roadmap' is not knowledge of what shipped."
Library at test timeprisma 7.10.0 (npm), verified 2026-09-05
Batteryprisma/v3-a · 0 tasks, 10 direct questions · probe window 6.7.0 to 7.9.0
Tool uses during test0 (a run with any tool use is void — we measure training knowledge, not retrieval)
Tested2026-09-05
Findings0, of which 0 chargeable

Findings

None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.

What it got right, and near misses

Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.

KindAPINote
contextprisma 6.7.0 LADDER GRADE — CORRECT. Named the Query Compiler Early Access, the queryCompiler and driverAdapters flags together, and the removal of the engine binary from the deployment, attributed to 6.7.0. Matches the release notes headline "Prisma ORM without Rust engines (Early Access)". Said PostgreSQL first where the notes say PostgreSQL and SQLite; not disqualifying. This is the highest real rung this arm graded CORRECT, so D = 6.7.0.
contextprisma 7.0.0 LADDER GRADE — ABSTAIN, and the sharpest abstention in the battery. "Cannot describe. I don't know whether this shipped, when, or what was in it." It then labelled the roadmap direction explicitly as "inference rather than knowledge". Both -cs arms, asked the identical question cold, described 7.0.0 correctly and dated it to November 2025. This arm had already placed its own boundary at 6.9.0 two questions earlier.
contextprisma 6.16.0, 7.4.0, 7.7.0 LADDER GRADES — ABSTAIN on all three. No content offered and none invented.
correctprisma 6.22.0 POISON RUNG. 6.22.0 does not exist — the last stable 6.x is 6.19.3 and no stable 6.20.0 or above was ever published. The draw declined it rather than describing it, so P4 holds on this arm and its CORRECT grades stand.
correctprisma 7.9.0 CEILING RUNG. 7.9.0 (2026-07-21) is above this subject's stated 2026-05 cutoff. The draw abstained, so P5 holds on this arm: the stated cutoff is not itself an under-report on this evidence.

Open questions from this run

Sources

Battery specification: prompts/prisma.md in the studio repo. Every finding above also carries its own citation.