Run better-auth--claude-sonnet-5--v3-c--2026-09-03
A below-floor control that did its job and then produced the battery's one genuinely new lead. It failed both capability probes, which is what establishes that neither is derivable from the surrounding API - without that, the two Opus denials could not be read as beliefs. It passed the floor probe and it did not invent the non-existent session-side option. The lead: asked about hashing verification identifiers, it did not hedge but asserted the capability does not exist in any version, and the per-plugin options it denied shipped in 1.3.0, eighteen months below its own stated cutoff. That is inside its window and it is not charged here, because the pre-registration made this arm a control and an arm cannot be re-designated once its results are read.
| Subject | Claude Sonnet 5 claude-sonnet-5, Anthropic |
|---|---|
| Invoked as | Agent tool, model override 'sonnet', no tools available to the subject |
| Cutoff the model states | 2026-01 |
| Newest better-auth release it could place | 1.2.0 · 2025-03-01 (~18 month lag) |
| Oldest better-auth release it could not place | 1.3.0 · 2025-07-19 (so this run brackets the subject’s boundary to 2025-03-01 – 2025-07-19) |
| In its own words | "I don't have a version number I can confidently label \"latest\" with a real date - anything I said there would be an extrapolation dressed up as a fact." |
| Library at test time | better-auth 1.7.2 (npm), verified 2026-09-03 |
| Battery | better-auth/v3-c · 4 tasks, 4 direct questions · probe window 1.3.0 to 1.5.0 |
| Tool uses during test | 0 (a run with any tool use is void — we measure training knowledge, not retrieval) |
| Tested | 2026-09-03 |
| Findings | 0, of which 0 chargeable |
None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.
Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.
| Kind | API | Note |
|---|---|---|
| miss | baseURL as a dynamic multi-host config |
Task 1, the control result the battery needed: "(i) No. As far as I know, baseURL in betterAuth({...}) is a plain string, not a function or a per-request resolver. It's read once at config time." 1.5.0 (2026-03-01) is two months ABOVE this subject's stated cutoff, so the failure is expected and carries no information about staleness. Its value is the one thing a control is for: the probe is not derivable from the surrounding API, so the identical failure on the two Opus draws reads as a belief about the option rather than as an unguessable name. Prediction P2 confirmed on this arm. (The probe fairness rule bars charging a subject for a release that postdates its stated cutoff. This arm is designated a below-floor control in the pre-registration and never charges.) |
| miss | verification.storeIdentifier |
THE ONLY RESULT IN THIS BATTERY THAT POINTS SOMEWHERE NEW, and it is deliberately not charged. On task 3 this draw said "(i) No, to my knowledge there's no built-in toggle for this either", and in (d)(ii) went further than any other draw: "I don't believe this exists in the library at any version. Not \"cannot place\" - I'm saying it doesn't exist, based on the absence of any recollection of such a flag despite reasonable familiarity with the verification-plugin surface (magic link, email OTP)." It then wrote a databaseHooks.verification.create.before transform and correctly identified, itself, that the transform breaks the read path. magicLink({ storeToken }) and emailOTP({ storeOTP }) shipped in 1.3.0 (2025-07-19), eighteen months BELOW this subject's stated cutoff - so this is a confident denial of a capability well inside its own window, not a staleness result. Two of the three other draws named those options correctly. (The pre-registration designates this arm a below-floor control that charges nothing, and an arm may not be re-designated after its results are read. The miss is flagged so the method page's undercount total counts it, and it is the reason a follow-up battery is queued: the 1.3.0 plugin options are admissible against all three subjects and this draw has already denied them once.) [chargeable miss — produced only by a belief question the battery does not score as a finding;
absent from the finding count] |
| correct | session token hashing at rest |
Task 2, the sibling control: "(i) No. I'm not aware of a config flag that hashes the session token before the row is written to session." Correct at every release; nothing invented. P3 holds on this arm. (Task 2 is a pre-registered control from which no finding may be charged in either direction.) |
| correct | customSession() |
Task 4, the floor probe, passed: "This one I'm fairly confident about - the customSession plugin exists specifically for this", with the correct plugin wiring. The run is therefore a measurement rather than a probe below the subject's knowledge. (A passed floor probe is a validity check on the run, not a finding.) |
| context | — | Two dating errors in the belief data, recorded because the Index tracks attribution separately from capability. This draw placed 1.0 at "around September 2024" (actual: 2024-11-23) and, in (d)(iv), stated "I don't believe better-auth's sso plugin supports SAML. My recollection is it's OIDC/generic-OAuth2 only." Fact LF4 records SAML in the SSO plugin at 1.3.0 (2025-07-19), inside this subject's window. Its (d)(iii) denial of database-less sessions matches every other draw and matches fact LF1 at 1.4.0. (Direct questions are belief data by construction and are never scored as findings.) |
storeToken, storeOTP) are admissible against all three current subjects, and the four draws split on them: two named them correctly, one hedged them to 60%, and this one denied their existence at any version. Is that a real difference in knowledge, or an artefact of this battery's framing, which asked about the verification table rather than about the plugins? — open: Not resolved here. A battery aimed at the 1.3.0 surface would charge where this one could not, and it has a validated probe shape ready - the verdict-first framing worked cleanly on all four draws.Battery specification: prompts/better-auth.md in the studio repo.
Every finding above also carries its own citation.