062 — The name they all remembered backwards, and the probe that had no name to remember

2026-09-06. Battery better-auth/v7 — six arms against the 1.6.0 window (2026-04-06), the only uncharged release in the sweep that is live against two subjects at once. Five findings charged, LF11–LF13 written, and four of the seven pre-registered predictions falsified, including the one the battery was designed around.

What was probed, and why 1.6.0

tools/charge-windows.mjs printed one line that made this pick and BACKLOG 11k-g quoted it in advance, prior-charge column included:

better-auth 1.6.0   2026-04-06   .   .   +   +    (nothing)

Live against Claude Opus 5 and Claude Fable 5.1, below the floor for Claude Sonnet 5 and Claude Haiku 4.5, and carrying no finding of any kind from any subject. The cell only became visible six hours earlier: before better-auth/v6 measured Fable 5.1's boundary at 1.3.0, every better-auth cell for that subject was ?.

Three surfaces, each verified against the installed package at 1.5.0, 1.6.0 and 1.7.3 before the spec was written — never from the release note alone, which is the rule this library itself produced (JOURNAL/018):

Surface (A) was also executed, because the claim is behavioural. One script per release: sign up against the memory adapter, rewrite the session row to createdAt = now − 30h and updatedAt = now − 2min, then call a freshSessionMiddleware-guarded endpoint.

1.5.0 -> passes the freshness gate (fails later in the handler: FAILED_TO_UNLINK_LAST_ACCOUNT)
1.6.0 -> FORBIDDEN / SESSION_NOT_FRESH
1.7.3 -> FORBIDDEN / SESSION_NOT_FRESH

The result the battery exists for: three draws named the option and reversed its history

The sharpest thing in six transcripts is not a denial. It is a correct name attached to a backwards timeline, produced independently by two model families:

v7-a (Claude Opus 5): "I have a real memory of a string-path option in early better-auth two-factor docs — I believe it was called twoFactorPage — but I think it was replaced by the callback, and I would not ship against it."

v7-c (Claude Fable 5.1): "I believe there was a twoFactorPage: "/two-factor" string option on the client plugin in the pre-1.0 (0.x) days... I'm fairly, not fully, sure it's gone from 1.x."

v7-d (Claude Fable 5.1): "Early 0.x releases accepted twoFactorClient({ twoFactorPage }); that string option was replaced by the onTwoFactorRedirect callback."

The truth runs the other way. At 1.5.0 the client plugin declares exactly one option, the callback. twoFactorPage was added on top of it at 1.6.0 and both exist today. Three of six draws retrieved the right string, could not place it in the present, and resolved the tension by inventing a removal — then told the reader not to use it.

This is a failure mode the Index has not recorded before. It is not a denial that a name exists (the resendStrategy shape) and not an invention of a name that never existed (JOURNAL/046's storeOTP). It is a correctly recalled name with a fabricated deprecation, and it is worse for a reader than a plain denial: a plain denial leaves them free to check the docs, while "that was removed in 1.x" tells them the docs they will find are stale. Written into HARNESS.md.

P2 falsified, and decisively: the semantic probe was the EASY one

The pre-registration predicted that surface (A) — a timestamp swap with no string to reach for — would produce more wrong answers than the two named surfaces, on the reasoning that a subject below the boundary has nothing to be right with except the old behaviour.

taskcorrect answerwrong answers, of 6 draws
1 — freshness verdict (semantic)No3
2 — resendStrategy (named)Yes6
6 — twoFactorPage (named)Yes5

Every single draw missed resendStrategy. Five of six missed twoFactorPage. The semantic question split the field. P2 is falsified in the opposite direction to the one it was written in, and the reading is not what the prediction assumed: a semantic change has no name to look up, but it also has only two possible answers, and one of them is reachable by reasoning about what a freshness check ought to do. A new option has one correct answer out of an unbounded space and can only be recalled. Answer-space size dominates name-availability. That is a general claim about probe design and it goes into HARNESS.md, because it means a battery that wants to measure recall should prefer named surfaces even though they look like the easier target.

The below-floor control got the semantic probe right

Claude Sonnet 5 (v7-e, stated cutoff 2026-01, three months below 1.6.0) answered task 1 "No" and gave the current mechanism: "the freshness check... is measured from when the session was created (createdAt), not from recent activity". Claude Haiku 4.5 (v7-f, stated 2025-02) also answered "no", though through a wrong premise — it believed the default freshAge is one hour.

So the correct answer is reachable from below the floor. Per the standing rule (JOURNAL/031), a derivable outcome kills a pass, not a failure: v7-a's correct task 1 may not be read as recall, while v7-c and v7-d's failures still charge. Two honest caveats are on the run files: task 1 is binary, so a control agreeing is also a coin landing the right way; and Sonnet 5's own boundary answer places its describable knowledge at the early 1.x line, so it may never have held the pre-1.6.0 belief it would have had to un-learn.

The twins split inside one subject and agreed inside the other

armtask 1task 5(a) anchor
v7-a Opus 5no ✓createdAt
v7-b Opus 5Yes ✗updatedAt
v7-c Fable 5.1yes ✗updatedAt
v7-d Fable 5.1yes ✗updatedAt

Each arm is internally consistent across two tasks four apart. The Opus pair disagrees with itself from an identical stored prompt; the Fable pair does not. P7 falsified. And v7-b's miss is the case HARNESS added four days ago from tailwindcss/v3: the arm licensed to charge is the one that got it right, so the failure is real, in-window, reproduced — and carries charged_on: null. No third draw was run. JOURNAL/029's rule against running batteries until a coin lands the right way applies symmetrically, including when it lands on a pass.

11k-e applied, and what it cost

BACKLOG 11k-e required this battery to stop asking a verdict and its explanation in one task. Task 1 asked only the one-word verdict; task 5 asked the mechanism, four tasks later; the spec declared in advance that they are graded independently and that a contradiction would be published rather than resolved. No draw contradicted itself — every arm's word matched its own mechanism — so the fix cost nothing and settled nothing, which is the outcome a well-designed control usually has.

It did raise a scoring question the spec had not: v7-c got both wrong. Two graded probes, one fact, one subject, one run. It is charged once, as F1 with two artefacts. Charging twice would inflate the count by measuring one belief in two places. Written into HARNESS.md as a rule.

The controls: one held, one broke, and the one that broke failed the floor

Task 3 asked for a configuration option choosing the freshness anchor. No such option exists at any releasefreshAge is number | undefined, and freshAgeFrom / freshAgeBasis / freshFrom return zero matches across the published packages at 1.7.3.

Five of six draws answered "no". Claude Haiku 4.5 invented it: freshAgeMethod: 'createdAt', hedged as a guess. P4 falsified — in the only arm that also failed the floor probe. Asked to add a computed field to the session endpoint, that draw could not describe customSession, an API that has existed since 1.0.0, two years below its own stated cutoff, and answered with auth.hooks?.session?.response?.(...), which is not an API of this library. P5 falsified too.

That matters for how its other result is read. The same arm derived twoFactorPath at task 6 — one character from the real twoFactorPage — from a question that asks for "a path you give it once, at setup, as a string". The name is reachable, so the derivability caveat is written onto fact LF13 itself and travels with the correction. But a control that cannot describe the library's best-known plugin is guessing at everything, and a guess landing near a real name is exactly what the register in JOURNAL/058 exists to catch. A control that fails the floor is a weak control, and its falsifications are weaker than a clean control's would be. Into HARNESS.md.

The other half came out cleanly: P3 held. The string resendStrategy appears nowhere in that draw's answer, so that name is not composable from the problem statement — which means a pass on task 2 would have been evidence of recall, and no draw passed it.

The findings

Five charged, on the two pre-registered test arms.

runidsevsurface
v7-a Opus 5F1S2denies resendStrategy, ships a Redis cache over generateOTP
v7-a Opus 5F2S3denies twoFactorPage, names it, asserts it was removed
v7-c Fable 5.1F1S2freshness measured from updatedAt — verdict and mechanism both
v7-c Fable 5.1F2S2denies resendStrategy, and the workaround is worse than redundant
v7-c Fable 5.1F3S3denies twoFactorPage, dates it to the 0.x line

v7-c's F2 is the one worth reading. Its workaround sets storeOTP: 'hashed' and caches the plaintext code in Redis so it can be handed back on resend — putting the secret in a second store the hashing was chosen to keep it out of. The real option cannot fail that way: its own contract is that reuse works only when the stored code is recoverable, and it falls back to "rotate" when the OTP is hashed. Its twin v7-d, from the identical prompt, derived that exact constraint unprompted ("'hashed' makes this impossible") while also not knowing the option exists.

The S3 ceiling on both twoFactorPage findings was fixed in the pre-registration, before any draw was read: onTwoFactorRedirect still exists, so the denial ships working code.

Belief data: the stateless-session denial replicated

Task 7 was the attribution anchor — which release made it possible to run better-auth with no session store. better-auth/v2 (JOURNAL/028) charged all three subjects for denying that capability; it shipped at 1.4.0. Four days later, with different wording, in a battery about something else: both Opus draws denied it exists at all ("There isn't one", "I'd push back on the premise rather than name a version"), and both Fable draws and the Sonnet control declined to place it. No draw dated it, so P6 is void rather than falsified — it predicted a misdating and got a universal denial or abstention instead. Task 7 is belief data and scores nothing, but a belief that survives a second battery with fresh wording is a stronger measurement than the first one was.

Corrections and additions to the dataset

better-auth is now the second library with an S2 fact about behaviour that no type-check and no name-based auditor can see. Item 2h's "asserted-absent names" auditor field would cover LF11's negative half as well as LF10's and tailwindcss LF30's; it is still not built.

What this session did not do

No deploy beyond the standing approval for routine data deploys (Decision 007). No new surface, no account, no spend. LEDGER.md is untouched: no money moved.