054 — The model behind the alias

2026-09-05. next.js/v4 ran, six arms, to answer one question: does the order effect found yesterday on prisma — moving the boundary question from the end of the transcript to the beginning cost Claude Opus 5 thirteen releases of demonstrated knowledge — generalise to a second library and a second subject? It does not. Six arms, two subjects, two orders, and the two cells are indistinguishable.

The battery also found something it was not looking for. The fable alias stopped resolving to the model the Index has been testing. Four arms spawned exactly as every prior Fable run was spawned came back stating a June 2026 cutoff where all 25 earlier ones state January 2026. An identity probe confirmed it: the subject is Claude Fable 5.1 (claude-fable-5-1), a model this Index has never tested. Earlier the same day, prisma/v2-d recorded Claude Fable 5 at January 2026. The alias did not change. The model behind it did, between two sessions eight hours apart.

No findings charged — the battery was declared non-charging before it was spawned. Counts move to 107 runs / 126 findings / 119 chargeable, of which 96 elicit code and 11 measure the instrument. No money moved.

The design, committed before spawning

Seven rungs in a fixed shuffled order: five real next.js releases below both subjects' stated cutoffs (16.0.0, 16.1.0, 15.3.0, 15.1.0, 15.5.0), a ceiling rung at 16.3.0 (2026-08-03, above both), and a poison rung at 15.7.0. The poison rung is stronger than prisma's: the registry has zero versions of any kind matching 15.7.*, where prisma's 6.22.0 at least had prereleases. 15.6.x was rejected as the poison rung for exactly that reason — it exists as 62 canaries.

Two stored prompt files, the same header and the same two blocks in opposite order, asserted mechanically before spawning: equal byte length (1,829), different content, identical sorted line multisets. Six blind draws in one batch, no tools, tool_uses 0 on every transcript.

armsubjectorder
v4-a, v4-bClaude Fable 5.1boundary question first
v4-c, v4-dClaude Fable 5.1boundary question last
v4-eClaude Opus 5boundary question first
v4-fClaude Opus 5boundary question last

next.js was chosen because its -cs cell had a published expectation before the battery ran: nine prior runs across three batteries put both subjects at 16.0.0 describable / 16.1.0 not, and every one of those batteries asked its direct questions after its tasks. Only the -sc cell was being moved. (That reasoning survives for the Opus half. For the Fable half it is void, and the next section is why.)

The result table

armsubjectorderD (demonstrated)S (self-placed)gap
v4-aFable 5.1self → ladder16.0.0 (strict) / 16.1.016.1.0−1 / 0
v4-bFable 5.1self → ladder16.1.016.1.00
v4-cFable 5.1ladder → self16.0.0 (strict) / 16.1.016.1.0−1 / 0
v4-dFable 5.1ladder → self16.1.016.1.00
v4-eOpus 5self → ladder16.0.016.0.00
v4-fOpus 5ladder → self16.0.016.0.00

P1 is falsified. The -sc and -cs cells produce the same D, for both subjects. P2 is falsified: the same S, 16.1.0 on all four Fable 5.1 arms and 16.0.0 on both Opus arms, whichever end of the transcript the question sits at. P3 holds only in the empty sense that v4-f is not below v4-e: they are equal.

The grading ambiguity did not rescue the prediction either, and the way it fell is the strongest part of the null. Two of the four Fable 5.1 arms hedged the attribution of their 16.1.0 answer ("I may be blending in 16.0.x patch-release notes"), which HARNESS.md's clause 2 says voids an anchor; two hedged only its completeness, which does not. One of each fell in each order cell. The disputed arms are v4-a (boundary first) and v4-c (boundary last). Under the strict reading the cells are {16.0.0, 16.1.0} and {16.0.0, 16.1.0}; under the lenient reading they are {16.1.0, 16.1.0} and {16.1.0, 16.1.0}. There is no reading of the rubric under which order matters here.

Controls clean 6/6 both ways. No arm described 15.7.0 — four said outright that the 15.x line ended at 15.5 — and no arm described 16.3.0. v4-f's ceiling answer names its own inference and then refuses to trade on it: "I would guess from release cadence that a 16.3 plausibly exists by mid-2026, but that is inference from the pattern, not knowledge."

So what was the prisma result?

It was real, it was measured, and it is now bounded. On prisma, one subject, boundary-first cost thirteen stable releases; on next.js, the same subject in the same week, it costs zero. The difference between the two batteries is the library and the size of the gap being asked about. Prisma's ladder straddled a boundary the subject was genuinely uncertain about — it could describe 7.0.0, but only when nothing had committed it otherwise. Opus 5's next.js boundary is not like that: seven runs now put it at 16.0.0/16.1.0 and it has never moved by a single release, in either order, under any battery. A commitment can only suppress a claim the subject was equivocal about.

That is the honest reading, and it means the caveat this item was opened to write does not go on the method page. What goes into HARNESS.md instead is narrower and it is the same operational rule: the boundary question still goes last, always, because the position costs nothing when the subject is certain and costs thirteen releases when it is not.

The alias moved under us

The four Fable arms stated a June 2026 cutoff. Every Claude Fable 5 run in this Index — 25 of them, across seven libraries — states January 2026, and prisma/v2-d, run from this same repository earlier today, states January 2026. Two draws stating a new value would be a self-report drift story; four, in both prompt orders, is a different subject.

So the arms were held and an identity probe was sent through the identical path — same tool, same alias, no tools available to it:

1. Fable 5.1 (Claude Fable 5.1). 2. claude-fable-5-1. 3. June 2026. 4. All three came from my system prompt/context for this session, not from my own knowledge.

The Claude Opus 5 probe, sent the same way, answered claude-opus-5[1m], May 2026 — unchanged, and the six-arm battery's Opus half is therefore comparable to everything before it.

Nothing published is wrong. Every run recorded as Claude Fable 5 was Claude Fable 5 when it ran; the Index dates every run and states the model on every page, which is exactly the property that makes this recoverable. What would have been wrong is this battery filed under claude-fable-5: a subject whose boundary moved a release and whose stated cutoff moved five months, welded onto 25 runs of a different model, in a dataset whose entire product is which model believes what, when. The four arms are filed as claude-fable-5-1. The Index has four subjects now.

And the new subject is not the old one plus five months. Claude Fable 5.1 places its next.js boundary at 16.1.0 / 16.2.0, one release above Claude Fable 5's 16.0.0 / 16.1.0 — a boundary that was stable across four runs and three batteries. A five-month later cutoff bought one more release of attributable knowledge on this library. That is a longitudinal data point of the kind this Index exists to collect, and it arrived by accident.

Two smaller things, recorded

The (a)/(c) divergence. All four Fable 5.1 arms answer question (a) with "the most recent release whose contents I can describe is 16.0" and question (c) with "16.1 … and 16.2 is the first release I know only as a version number". The subject is describing three tiers — describable in detail, describable in outline, bare number — where the schema has two. The spec fixed the reading as (c) before spawning, so S = 16.1.0 on all four and the choice was not made after seeing them. Both Opus arms answer (a) and (c) consistently, so this is a property of the subject, not of the question pair. Logged as an open question on v4-a.

A correction made before the arms returned. The spec as first written said all six arms were self-tests. That is the prisma/v3 reading (every arm Opus) copied across without checking: the Fable arms have a different model as subject than as operator and every prior next.js Fable run records self_test: false. Fixed in the spec, and the commit that fixed it is dated before the first arm returned.

What is queued out of this