Run better-auth--claude-fable-5--v3-d--2026-09-03
The second below-floor control, and the draw that settled the battery's other half. It failed the baseURL probe like every other arm, which is what licenses reading the two Opus failures as beliefs rather than as noise. But on verification hashing it answered yes and named storeOTP and storeToken with their correct value sets - from eighteen months below the target release - which is what proved the Index's own newly written fact LF9 had claimed too much. It also gave the most accurate version attribution of the four draws, placing SAML SSO exactly at 1.3.0. A control arm that corrects the operator is a better outcome than one that merely confirms a floor.
| Subject | Claude Fable 5 claude-fable-5, Anthropic |
|---|---|
| Invoked as | Agent tool, model override 'fable', no tools available to the subject |
| Cutoff the model states | 2026-01 |
| Newest better-auth release it could place | 1.3.0 · 2025-07-19 (~18 month lag) |
| Oldest better-auth release it could not place | 1.4.0 · 2025-11-22 (so this run brackets the subject’s boundary to 2025-07-19 – 2025-11-22) |
| In its own words | "The most recent release whose contents I can actually describe is 1.3, around July 2025. I have a vague sense that development continued after that - patch releases on 1.3.x through late 2025, and possibly a 1.4 - but I cannot describe what's in anything past 1.3 with confidence." |
| Library at test time | better-auth 1.7.2 (npm), verified 2026-09-03 |
| Battery | better-auth/v3-d · 4 tasks, 4 direct questions · probe window 1.3.0 to 1.5.0 |
| Tool uses during test | 0 (a run with any tool use is void — we measure training knowledge, not retrieval) |
| Tested | 2026-09-03 |
| Findings | 0, of which 0 chargeable |
None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.
Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.
| Kind | API | Note |
|---|---|---|
| miss | baseURL as a dynamic multi-host config |
Task 1, the second control result: "(i) No. baseURL is typed as a plain optional string. There is no function form, no array form." It then proposed dropping baseURL entirely so the library infers the origin from the request, gated by trustedOrigins - a different workaround from the two Opus draws and, notably, the closest of the four to the shape of the real feature without ever reaching it. 1.5.0 is two months above this subject's stated cutoff. Prediction P2 confirmed on this arm too: neither Opus denial can be an artefact of an unguessable name. (The probe fairness rule bars charging a release that postdates the subject's stated cutoff. This arm is a designated below-floor control.) |
| correct | verification.storeIdentifier |
THE CONTROL THAT ANSWERED THE PROBE THE TEST ARM WAS SUPPOSED TO FAIL. On task 3 this below-floor draw answered "(i) Yes - but per plugin, not as one global switch on the verification table. The email-OTP plugin has a storeOTP option and the magic-link plugin grew an equivalent storeToken; both accept \"plain\" (default), \"hashed\", and (for OTP at least) \"encrypted\", plus a custom-hasher form." Every clause of that is correct against the published 1.3.0 declarations, including that encrypted exists for OTP and not for magic-link. Its (iii) answer on deterministic hashing is correct. Per JOURNAL/030 the outcome of this probe is therefore DERIVABLE-at-1.3.0 for the plugin half - which is not a disqualification of anything, because the plugin half is genuinely old knowledge; it is the evidence that made the Index correct its own fact LF9 the same session. (Correct answers are not stale priors. Recorded because a below-floor control passing a probe is the outcome that kills a probe's reading, and here it killed the Index's own claim rather than a subject's.) |
| correct | session token hashing at rest |
Task 2, the sibling control: "(i) No. To my knowledge there is no config that hashes the session.token column before write", plus the unprompted and correct observation that a databaseHooks.session.create hook cannot fix it because the library's own lookup would then miss. Nothing invented. P3 holds on all four arms. (Task 2 is a pre-registered control from which no finding may be charged in either direction.) |
| correct | customSession() |
Task 4, the floor probe, passed: customSession() plus customSessionClient<typeof auth>(), with the correct note that it runs on every getSession and that a stored rather than computed field belongs in additionalFields. Floor confirmed on all four arms. (A passed floor probe is a validity check on the run, not a finding.) |
| context | — | The most accurate belief data of the four draws. It placed SAML SSO at "1.3, ~July 2025" - fact LF4 records 1.3.0, 2025-07-19, exactly right - and placed the verification-hashing options in "the 1.2.x line, roughly mid-2025 ... around 1.2.7", which is one minor early against the real 1.3.0 but is the only draw to give a number at all. Its (d)(iii) denial of database-less operation matches every other draw against fact LF1 (1.4.0), which is two months above its stated cutoff and therefore not chargeable here. (Direct questions are belief data by construction and are never scored as findings.) |
Battery specification: prompts/better-auth.md in the studio repo.
Every finding above also carries its own citation.