127 — The two seats that share one grader

2026-09-10 · distribution lane · DISTRIBUTION D1 / BM2 step 1

BM2's fourth Class A seat is filled and admitted. A4's drawn fact, LF15, was refused this morning on the two-boundary rule (JOURNAL/125), and the seat takes the next S1 reserve entry in task order: LF1, "constructing the client with no arguments is no longer supported; Prisma 7 requires either a driver adapter or an accelerateUrl". The gate admits it on the same twenty-three installed prisma releases A1, A2 and A3 were measured on:

A4   A  LF1   falling   @7.0.0
     correct  +++++++++++++++++++++++
     stale    ++++++++++++...........

The stale artifact is LF1's own stale_codenew PrismaClient(), no adapter, no connection string in code. The half of that row that had to be measured is the left half, not the right one. The artifact passes below 7.0.0 only because the harness's schema carries a connection string on those rungs: spell() injects a literal url into every 6.x datasource block, which is what a reader on that release had to write too. So the pass below the boundary is the fact's own claim — the client reads its datasource out of the schema — rather than an artefact of a schema that happens to be permissive. Had the 6.x schema been urlless, the artifact would have failed everywhere and the seat would have looked like the eighth amendment's identical-rows refusal for a reason that was ours.

What the seat forced into the open, and it is about the SET rather than the task

A3 grades LF3 (datasourceUrl). A4 grades LF1. Both facts are dated prisma 7.0.0, both are the constructor refusing what it used to accept, and at the graded release they have one correct answer: pass a driver adapter. So A4's acceptance assertion is A3's — literally, one function called with a different fact id — and the two reference artifacts are byte-identical.

That is written into tools/benchmark/tasks/prisma.mjs as one constructsClient(factId) rather than as two functions with the same body, because the alternative is two copies that can drift apart while looking like independent instruments. An assertion that could tell a correct A4 answer from a correct A3 answer would have to grade the source text, which §9 forbids everywhere else in this benchmark.

The claim this entry retracts is its own, and it was retracted before it was committed

The first draft of that comment said the two seats "cannot vary independently". The cross-measurement was run to quote it rather than to test it, and it says otherwise:

                       A3's assertion        A4's assertion
  6.19.0  A3.stale     PASS                  PASS
          A4.stale     PASS                  PASS
  7.10.0  A3.stale     FAIL                  FAIL
          A4.stale     FAIL                  FAIL

Two things fall out of the four failures at 7.10.0.

The graders separate the two stale beliefs not at all. Either stale answer fails either seat. If the seats were graded on the artifacts alone, they would be one measurement.

The cells are still not one cell, and A3's own committed measurement is the counter-example. { adapter, datasourceUrl } satisfies LF1 — an adapter is passed — and still holds LF3's belief, and it fails at 7.0.0 and 7.10.0 on datasourceUrl alone. A subject that writes that for A3's prompt and { adapter } for A4's fails A3 and passes A4. The independence between these cells lives entirely in what each prompt draws out of a subject: A3's asks for a client "configured to use that connection string", which is exactly what datasourceUrl was the shorthand for, and A4's names no override at all, which is the case LF1's stale_belief describes. Which is a property of the prompts and therefore something the arms measure, not something this file may predict.

And a third thing, which is not about verdicts. The four failures print two different sentencesUnknown property datasourceUrl provided to PrismaClient constructor. against A3's artifact and PrismaClient was instantiated without any options. A driver adapter is required to connect to your database. against A4's. §9 requires a failing round to hand back the toolchain's own words, so the seats' retry channels differ even where their verdicts cannot, and rounds_to_pass — one of the numbers BM2 publishes — can differ between them for reasons that belong to the subject.

What it costs the results, fixed before any arm has run

Property 1 of the BM2 draw already forbids any interval, significance claim or "N of 12" figure whose derivation assumes the twelve tasks are independent draws, on the ground that eleven of them probe one release. This is that constraint one level sharper and at seat level: A3 and A4 — and A7, if LF2 is authored into it — are one instrument, and any published count over BM2's tasks has to name the group.

The rule is deliberately not "drop the duplicate seats". The reserve order was fixed by draw.mjs before any of this was known, and refusing a seat because a seat authored earlier resembles it is discovery order deciding the set — the property §7 exists to remove, and the same trap the substitution order avoided this morning. Three of BM2's twelve seats descending from one removal is a fact about prisma 7 and about the draw, and the honest treatment is to publish it beside the counts, not to quietly re-draw until the set looks more varied than the library is.

Nothing else moved

A1, A2 and A3 re-gated on the same ladder after the assertion file was refactored: all admitted, every profile byte-identical to the run captured before the first edit. Gate selftest 46, runner selftest 42, MCP selftest 54, all five --check builds green. No dataset count moved — admitting a benchmark task publishes no finding — and the counts stand where JOURNAL/125 left them: 161 runs, 167 findings (160 chargeable), 8 libraries.

What is left of BM2: eight drawn Class A assertions plus A7's substitute (LF2), the six Class N tasks, BM2's predictions, then the arms.