What Claude Sonnet 5 gets right about better-auth — battery v3-c, tested 2026-09-03

Run better-auth--claude-sonnet-5--v3-c--2026-09-03

Summary

A below-floor control that did its job and then produced the battery's one genuinely new lead. It failed both capability probes, which is what establishes that neither is derivable from the surrounding API - without that, the two Opus denials could not be read as beliefs. It passed the floor probe and it did not invent the non-existent session-side option. The lead: asked about hashing verification identifiers, it did not hedge but asserted the capability does not exist in any version, and the per-plugin options it denied shipped in 1.3.0, eighteen months below its own stated cutoff. That is inside its window and it is not charged here, because the pre-registration made this arm a control and an arm cannot be re-designated once its results are read.

SubjectClaude Sonnet 5 claude-sonnet-5, Anthropic
Invoked asAgent tool, model override 'sonnet', no tools available to the subject
Cutoff the model states2026-01
Newest better-auth release it could place1.2.0 · 2025-03-01 (~18 month lag)
Oldest better-auth release it could not place1.3.0 · 2025-07-19 (so this run brackets the subject’s boundary to 2025-03-01 – 2025-07-19)
In its own words"I don't have a version number I can confidently label \"latest\" with a real date - anything I said there would be an extrapolation dressed up as a fact."
Library at test timebetter-auth 1.7.2 (npm), verified 2026-09-03
Batterybetter-auth/v3-c · 4 tasks, 4 direct questions · probe window 1.3.0 to 1.5.0
Tool uses during test0 (a run with any tool use is void — we measure training knowledge, not retrieval)
Tested2026-09-03
Findings0, of which 0 chargeable

Findings

None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.

What it got right, and near misses

Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.

KindAPINote
missbaseURL as a dynamic multi-host config Task 1, the control result the battery needed: "(i) No. As far as I know, baseURL in betterAuth({...}) is a plain string, not a function or a per-request resolver. It's read once at config time." 1.5.0 (2026-03-01) is two months ABOVE this subject's stated cutoff, so the failure is expected and carries no information about staleness. Its value is the one thing a control is for: the probe is not derivable from the surrounding API, so the identical failure on the two Opus draws reads as a belief about the option rather than as an unguessable name. Prediction P2 confirmed on this arm. (The probe fairness rule bars charging a subject for a release that postdates its stated cutoff. This arm is designated a below-floor control in the pre-registration and never charges.)
missverification.storeIdentifier THE ONLY RESULT IN THIS BATTERY THAT POINTS SOMEWHERE NEW, and it is deliberately not charged. On task 3 this draw said "(i) No, to my knowledge there's no built-in toggle for this either", and in (d)(ii) went further than any other draw: "I don't believe this exists in the library at any version. Not \"cannot place\" - I'm saying it doesn't exist, based on the absence of any recollection of such a flag despite reasonable familiarity with the verification-plugin surface (magic link, email OTP)." It then wrote a databaseHooks.verification.create.before transform and correctly identified, itself, that the transform breaks the read path. magicLink({ storeToken }) and emailOTP({ storeOTP }) shipped in 1.3.0 (2025-07-19), eighteen months BELOW this subject's stated cutoff - so this is a confident denial of a capability well inside its own window, not a staleness result. Two of the three other draws named those options correctly. (The pre-registration designates this arm a below-floor control that charges nothing, and an arm may not be re-designated after its results are read. The miss is flagged so the method page's undercount total counts it, and it is the reason a follow-up battery is queued: the 1.3.0 plugin options are admissible against all three subjects and this draw has already denied them once.) [chargeable miss — produced only by a belief question the battery does not score as a finding; absent from the finding count]
correctsession token hashing at rest Task 2, the sibling control: "(i) No. I'm not aware of a config flag that hashes the session token before the row is written to session." Correct at every release; nothing invented. P3 holds on this arm. (Task 2 is a pre-registered control from which no finding may be charged in either direction.)
correctcustomSession() Task 4, the floor probe, passed: "This one I'm fairly confident about - the customSession plugin exists specifically for this", with the correct plugin wiring. The run is therefore a measurement rather than a probe below the subject's knowledge. (A passed floor probe is a validity check on the run, not a finding.)
context Two dating errors in the belief data, recorded because the Index tracks attribution separately from capability. This draw placed 1.0 at "around September 2024" (actual: 2024-11-23) and, in (d)(iv), stated "I don't believe better-auth's sso plugin supports SAML. My recollection is it's OIDC/generic-OAuth2 only." Fact LF4 records SAML in the SSO plugin at 1.3.0 (2025-07-19), inside this subject's window. Its (d)(iii) denial of database-less sessions matches every other draw and matches fact LF1 at 1.4.0. (Direct questions are belief data by construction and are never scored as findings.)

Open questions from this run

Sources

Battery specification: prompts/better-auth.md in the studio repo. Every finding above also carries its own citation.