What Claude Sonnet 5 gets right about better-auth — battery v5-c, tested 2026-09-03

Run better-auth--claude-sonnet-5--v5-c--2026-09-03

Summary

The control arm that did not hold, and it is this battery's most consequential result. Pre-registered as a control because both of this subject's better-auth/v4 draws denied the phone-number storage option correctly, it inverted: yes, storeOTP, shipped in the config as the fix. P2 falsified. That does not touch F1 - the charge on v5-a rests on non-compiling code, not on a contrast between arms - but it does change what the finding is about. Across the two batteries this belief is Opus 5 four draws for four, and Sonnet 5 one for three: the invention is not the property of a single subject that the pre-registration assumed, and the wording between v4 and v5 is the variable that moved. This draw also denied a real option on the one-time-token plugin, which the other three test draws all named.

SubjectClaude Sonnet 5 claude-sonnet-5, Anthropic
Invoked asAgent tool, model override 'sonnet', no tools available to the subject
Cutoff the model states2026-01
Newest better-auth release it could place1.1.0 · 2024-12-01 (~13 month lag)
Oldest better-auth release it could not place1.2.0 · 2025-03-01 (so this run brackets the subject’s boundary to 2024-12-01 – 2025-03-01)
In its own words"I have a vague, unreliable sense of version numbers in the 1.2-1.3 range existing (very roughly mid-to-late 2025), but that's closer to name-recognition than knowledge. The most recent release whose actual contents I could describe with any real confidence is more like the 1.1.x line."
Library at test timebetter-auth 1.7.2 (npm), verified 2026-09-03
Batterybetter-auth/v5-c · 3 tasks, 4 direct questions · probe window 1.3.0 to 1.7.2
Tool uses during test0 (a run with any tool use is void — we measure training knowledge, not retrieval)
Tested2026-09-03
Findings0, of which 0 chargeable

Findings

None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.

What it got right, and near misses

Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.

KindAPINote
missphoneNumber({ storeOTP }) Task 1, the probe, and the result that falsifies this battery's P2. This subject denied the same option correctly on BOTH of its better-auth/v4 draws, four hours earlier, from a prompt asking about the same plugin and the same surface. Here it answered "(i) Yes. My recollection is that the phoneNumber plugin does accept an option that changes what gets persisted for the OTP", named storeOTP at "medium-high confidence", and shipped it with the comment "<- don't persist the raw code". phoneNumber({ storeOTP }) is a TS2353 at 1.3.0, 1.5.0 and 1.7.2; the draw's entire remaining configuration - sendOTP, otpLength, expiresIn, allowedAttempts - is real and compiles. (This arm was pre-registered as a control, on the ground that all four non-Opus draws of better-auth/v4 denied this option correctly. It did not hold, and the pre-registration is not revised after the fact - an arm may not be re-designated after its results are read (JOURNAL/044). Counted in the method page's undercount total. Whether it is worth a battery of its own is doubtful and BACKLOG says why: three draws of this subject on this surface now read deny / deny / invent, which is intermittent, and JOURNAL/029 already ruled out running batteries until a coin lands in the charging arm.) [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing; absent from the finding count]
missoneTimeToken storeToken Task 2, the exists-control: denied oneTimeToken({ storeToken }) - "one-time-token plugin: no (low confidence guess, not a confirmed fact)" - and restated it in (d)(ii) as "I don't believe this exists as a named plugin option at all". It shipped in 1.3.0 and type-checks clean at 1.3.0, 1.5.0 and 1.7.2. The other three test draws all named it correctly, two of them with its exact { type: 'custom-hasher', hash } form. (Task 2 is a pre-registered control from which no finding may be charged in either direction, and this arm is a pre-registered control besides. Two independent bars, either of which is sufficient.) [chargeable miss — produced only by a belief question the battery does not score as a finding; absent from the finding count]
correcttwoFactor otpOptions.storeOTP Task 2, the other half: named twoFactor({ otpOptions: { storeOTP } }) correctly, at the correct nesting, though on stated "medium confidence" and by explicit analogy to the phone-number option it had just invented. The right answer reached through a wrong premise - the analogy runs from a plugin that does not have the option to one that does.
correctcustomSession Task 3, the floor probe: passed. customSession named as the mechanism, correctly distinguished from session.additionalFields as "a different tool for a different job" - persisted columns versus a computed wrapper - on stated medium confidence.

Sources

Battery specification: prompts/better-auth.md in the studio repo. Every finding above also carries its own citation.