Run prisma--claude-opus-5--v2-b--2026-09-05 · self-test: the subject is the operator
Blind twin of prisma/v2-a, charging nothing by design. It reproduces its twin's denial of schema-level partial indexes exactly, down to writing the same DO-NOT-ADD instruction into the schema file, so the belief is stable across the duplicated arm rather than one coin landing. It also reproduces the anchor miss, placing 7.4.0's query-plan caching at 6.7/7.0. Where the two twins diverge is the boundary: this draw described 7.0.0's contents correctly where v2-a called 7.0.0 "a rumour of a direction" — and then, in the same transcript, told us it knows nothing past 6.9/6.10. Its own self-report undercuts its own demonstration.
| Subject | Claude Opus 5 claude-opus-5, Anthropic |
|---|---|
| Invoked as | Agent tool, model override 'opus', no tools available to the subject |
| Cutoff the model states | 2026-05 |
| Newest prisma release it could place | 7.0.0 · 2025-11-19 (~5 month lag) |
| Oldest prisma release it could not place | 7.1.0 · 2025-12-03 (so this run brackets the subject’s boundary to 2025-11-19 – 2025-12-03) |
| In its own words | "The most recent version number I have any belief about is somewhere in the Prisma 7.x line — plausibly around 7.3-7.5, early-to-mid 2026. I hold that loosely... My belief about 'latest' is based purely on extrapolating Prisma's roughly-monthly minor cadence from what's in my training data." |
| Library at test time | prisma 7.10.0 (npm), verified 2026-09-05 |
| Battery | prisma/v2-b · 3 tasks, 4 direct questions · probe window 7.4.0 to 7.10.0 |
| Tool uses during test | 0 (a run with any tool use is void — we measure training knowledge, not retrieval) |
| Tested | 2026-09-05 |
| Findings | 0, of which 0 chargeable |
None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.
Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.
| Kind | API | Note |
|---|---|---|
| miss | @@index([...], where: ...) / @@unique([...], where: ...) |
The same failure v2-a charges as F1, reproduced on the blind twin. Verdict-first, task 1(i): "No." In prose: "Prisma's schema language has no way to attach a WHERE clause to @@index or @@unique. This has to be a hand-written migration, and the schema file has to be kept deliberately silent about it." Restated in (d)(i): "does not exist. Prisma has never supported partial indexes in the schema, to my knowledge. It's one of the longest-running open feature requests." Its schema block carries the same durable instruction its twin's does — "DO NOT add @unique to email and DO NOT add an @@index([email])... Prisma cannot express WHERE \"deletedAt\" IS NULL" — and it reasons its way to the same lost findUnique. Charged on v2-a, not here. (The -b draw of a duplicated test arm charges nothing (JOURNAL/028). The failure IS counted, once, on v2-a. Recorded here so the agreement between the twins is visible: on this surface the two draws agree exactly, which is the opposite of what they do on the boundary question.) [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing;
charged as a finding on prisma--claude-opus-5--v2-a--2026-09-05] |
| correct | @@index([...], include: [...]) |
Task 2, the covering-index probe, answered CORRECTLY. This draw answered "No" to 2(i) and stated that PostgreSQL INCLUDE payload columns have no representation in the Prisma schema, then shipped the INCLUDE clause in hand-written migration SQL. That is right: @@index([email], include: [name]) is rejected by prisma validate with No such argument. at 7.4.0 and at 7.10.0 (fact LF27). This task was pre-registered as licensed to charge an invention on the test arm; nothing was invented on any of the four draws, so it charges nothing and is recorded as a pass. (The answer is correct against the shipped validator. Prediction P3 — that at least one draw would over-extend 7.4.0's new index argument into include: — is FALSIFIED 0 of 4.) |
| correct | @@index([...], sort / map) |
Task 3, the floor probe, PASSED. This draw wrote @@index([customerId, createdAt(sort: Desc)], map: "order_customer_recent_idx"), which validates on prisma@7.10.0. Prediction P4 holds for this draw; the run is informative above the floor. (Correct code on the current release.) |
| miss | query plan cache |
THE ANCHOR, P5 falsified on this draw too. (d)(iii) answered "Prisma 6.7 (~May 2025), preview, as queryCompiler; default in 7.0", with the mechanism described correctly — "query compilation moves into TypeScript and compiled plans are cached per query shape" — and the release eight months early. Both Opus 5 draws made the same substitution independently, which is what makes it worth a note rather than a shrug. (Direct questions are belief data and are never scored as findings.) |
| context | — | BOUNDARY, and an internal contradiction inside one draw. This draw described 7.0.0's contents correctly and in the right terms — "the Rust-free query engine (query compiler) becomes the default; the new prisma-client generator replaces prisma-client-js as the default, generating ESM output into your source tree rather than node_modules" — which is the attribution prisma/v1 and v1r-a also produced, and which places its boundary at 7.0.0. But asked directly which release it knows only as a number, the SAME draw answered "roughly 6.9/6.10 (June 2025)", four releases below the one it had just described. Its own two answers to the boundary question are inconsistent, in one transcript, without the prompt changing. Recorded as 7.0.0 on the field's definition — newest release whose contents were correctly attributed — with the contradiction carried here rather than resolved silently. (Belief data, and an observation about the instrument.) |
| context | — | The pre-registered ordering effect did not fire, and the direction it would have pushed in is worth recording. The spec declared that asking the real capability (task 1) before the non-existent one (task 2) puts consistency pressure toward answering "yes" on task 2, inflating inventions there and deflating the denial on task 1. Every draw answered "no" to both, so no such pressure is visible. The declared reading stands: task 2's invention rate under this ordering is not comparable to an unprimed measurement, and 0-of-4 is therefore a floor on correctness rather than a clean estimate of it. (Instrument data.) |
| context | — | ERRATUM against this battery's own pre-registration, recorded rather than quietly dropped. The four-cell reading table in prompts/prisma.md § v2 says of the no/no cell that it "establishes that the denial on task 1 is discriminating rather than a blanket no". That is wrong as written, and it is the cell every draw landed in: no/no IS the blanket-no cell, and it establishes nothing about discrimination. Only the yes/no cell does. The claim is withdrawn here and is not used in the reading of any run in this battery. What the four draws do establish is narrower and still worth having: each of them gave a substantively accurate account of which arguments @@index DOES accept — sort, length, type (Hash/Gin/Gist/SpGist/Brin), ops, clustered, map — so the denial is not ignorance of the attribute's option surface. It is an option surface that is accurate as of 4.0.0 and closed to additions after it. (A correction to the battery spec, not an observation about the subject.) |
Battery specification: prompts/prisma.md in the studio repo.
Every finding above also carries its own citation.