Run prisma--claude-opus-5--v3-e--2026-09-05 · self-test: the subject is the operator
The mention-only control, single by design. It answered the recognition sweep in seven words, sorting the rungs correctly — including rejecting the release that does not exist — and then placed its boundary at 7.0.0, matching the highest -cs arm and exceeding both -sc arms. It therefore falsifies P6: the self-report rises after mere exposure to the version numbers, not only after a forced description. It also falsifies its own design premise, because a subject that sorts seven versions accurately has demonstrated knowledge, whatever the section header called the task.
| Subject | Claude Opus 5 claude-opus-5, Anthropic |
|---|---|
| Invoked as | Agent tool, model override 'opus', no tools available to the subject; prompt sent verbatim from prompts/sent/prisma-v3-mo.txt |
| Cutoff the model states | 2026-05 |
| Newest prisma release it could place | 7.0.0 · 2025-11-19 (~6 month lag) |
| Oldest prisma release it could not place | 7.1.0 · 2025-12-03 (so this run brackets the subject’s boundary to 2025-11-19 – 2025-12-03) |
| In its own words | "So '7.x is latest' is an inference about a moving target, not a fact I hold." |
| Library at test time | prisma 7.10.0 (npm), verified 2026-09-05 |
| Battery | prisma/v3-e · 0 tasks, 10 direct questions · probe window 6.7.0 to 7.9.0 |
| Tool uses during test | 0 (a run with any tool use is void — we measure training knowledge, not retrieval) |
| Tested | 2026-09-05 |
| Findings | 0, of which 0 chargeable |
None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.
Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.
| Kind | API | Note |
|---|---|---|
| context | recognition sweep |
CONTROL ARM, one word per rung and nothing else, exactly as instructed. Recognised: 7.0.0 and 6.7.0. Unrecognised: 6.22.0, 7.7.0, 7.9.0, 7.4.0, 6.16.0. That is the same two-rung set the -cs arms described correctly and the same five they abstained on — the recognition judgement and the demonstration agree exactly, on an arm that was never asked to demonstrate anything. |
| context | P6 |
P6 FALSIFIED, and the design premise behind it with it. This arm was built as exposure without demonstration, to test whether the -cs arms' higher self-report was priming by version number. Its S is 7.0.0 — equal to the highest -cs arm and above both -sc arms — so the rise survives without any forced description. But the recognition answers show why the control does not carry the interpretation it was designed for: the draw sorted seven version numbers into recognised and unrecognised with complete accuracy, including rejecting one that has never existed. Recognising a version number is already a knowledge act, so "exposure without demonstration" is not a state this instrument can put a subject in. |
| correct | prisma 6.22.0 |
POISON RUNG. 6.22.0 does not exist — the last stable 6.x is 6.19.3 and no stable 6.20.0 or above was ever published. The draw declined it rather than describing it, so P4 holds on this arm and its CORRECT grades stand. |
| correct | prisma 7.9.0 |
CEILING RUNG. 7.9.0 (2026-07-21) is above this subject's stated 2026-05 cutoff. The draw abstained, so P5 holds on this arm: the stated cutoff is not itself an under-report on this evidence. |
Battery specification: prompts/prisma.md in the studio repo.
Every finding above also carries its own citation.