What Claude Haiku 4.5 gets right about better-auth — battery v7-f, tested 2026-09-06

Run better-auth--claude-haiku-4-5--v7-f--2026-09-06 · self-test: the subject is the operator

Summary

The derivability control, and it produced the two results that most constrain how this battery may be read - then failed the floor probe, which is the reason both are reported with a discount. It failed task 4, unable to describe customSession, an API two years below its own cutoff, so it is a weak control rather than a clean one. It invented the non-existent freshness-anchor option at task 3 (freshAgeMethod), falsifying P4 in the only arm that failed the floor. And it derived twoFactorPath at task 6 - one character from the real twoFactorPage - which is the outcome that discounts any pass on that surface without excusing the denials charged against it. On the other side, P3 held: resendStrategy appears nowhere in its answer, so that name is not composable from the question. Charges nothing, and never could.

SubjectClaude Haiku 4.5 claude-haiku-4-5, Anthropic
Invoked asAgent tool, model alias "haiku"; prompt sent verbatim from prompts/sent/better-auth-v7.txt, no tools available to the subject
Cutoff the model states2025-02
Newest better-auth release it could placenot established in this run
In its own words"I believe better-auth is around v0.11–v0.12, but I'm not certain... I don't have confident knowledge of specific releases and their detailed changelogs as of February 2025."
Library at test timebetter-auth 1.7.3 (npm), verified 2026-09-06
Batterybetter-auth/v7-f · 7 tasks, 3 direct questions · probe window 1.6.0 to 1.7.3
Tool uses during test0 (a run with any tool use is void — we measure training knowledge, not retrieval)
Tested2026-09-06
Findings0, of which 0 chargeable

Findings

None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.

What it got right, and near misses

Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.

KindAPINote
contextcustomSession Task 4, the floor probe, FAILED, and that is the first thing to read about this arm. Asked to add a computed field to the session endpoint it answered "I'm uncertain of the exact mechanism" and produced auth.hooks?.session?.response?.(...), which is not an API of this library. customSession has existed since 1.0.0, two years below this subject's cutoff. P5 predicted all six draws would pass the floor and this draw falsified it. Consequence, and it is not cosmetic: everything else this arm says about better-auth is uninformative as a control, because a control's value is that it knows the library and not the release. The two results below are reported with that discount attached rather than treated as measurements.
contextsession.freshAge anchor option (does not exist) Task 3, the pre-registered non-existent control, invented: "yes", followed by session: { freshAge: 60 * 60, freshAgeMethod: 'createdAt' } with "freshAgeMethod is my best guess, but it could differ". No such option exists at any release, under that spelling or any other - freshAgeFrom, freshAgeBasis and freshFrom return zero matches across the published packages at 1.7.3. This falsifies P4, which predicted no draw would claim the option exists. It falsifies it only in the arm that failed the floor: all four above-floor draws and the below-floor Sonnet 5 control answered "no". Recorded as a chargeable miss that this battery may not charge - task 3 is pre-registered as a control and no finding may be drawn from it in either direction (the same bind JOURNAL/045 was in). (Task 3 was pre-registered as a control and a task may not be re-designated after its results are read (JOURNAL/045). The arm is also a control that never charges. Counted in the method page's undercount total.) [chargeable miss — produced only by a belief question the battery does not score as a finding; absent from the finding count]
contexttwoFactorClient({ twoFactorPage }) Task 6, the derivability result this battery needed, and it lands against us. "yes", followed by twoFactorClient({ twoFactorPath: '/auth/two-factor' }) - the right shape and one word off the right name, from a subject two years below the release that added it, while stating "the exact option name is uncertain". The name is reachable from the question, which asks for "a path you give it once, at setup, as a string". Per JOURNAL/031 the rule is stated in advance and applies as written: a derivable outcome kills a pass, not a failure. The three denials charged against twoFactorPage (F2 on v7-a, F3 on v7-c) stand; what nobody may now claim is that a correct answer on this surface demonstrates recall. The caveat is written onto fact LF13 itself so it travels with the correction. Discount this result further for the floor failure above: a draw that cannot describe customSession is guessing at everything, and one of its guesses landing near a real name is what P3 was designed to detect.
contextemailOTP({ resendStrategy }) Task 2, and the half of the derivability question that came out the other way: "no", and the string resendStrategy appears nowhere in the answer. It declined to write a configuration at all - "I cannot confidently write the config without being unsure of the actual API." P3 holds. The name is not composable from the problem statement, so a pass on task 2 would have been evidence of recall - and no draw in the battery passed it. (Below the floor by fourteen months, and a control arm that never charges.)
contextsession.freshAge (measured from session.createdAt) Task 1: "no" - the correct verdict, reached through a wrong premise. "The default freshAge in better-auth is typically 1 hour (measuring from createdAt). A session created 30 hours ago exceeds this." The default is 60 * 60 * 24, not one hour, at every release in and below the window. This is the register the JOURNAL/058 rule asks for: the guessing control can guess and hit, and the way to tell is to read what it got right on the way. Here it got the anchor right and the constant wrong, which is a pattern of guessing, not of knowing. Task 5(b) missed as well - it offered freshAge: Infinity or requireFreshSession: false, neither of which is the off switch.
contextknowledge boundary Task 7 and the direct questions: no answer of any kind. It could not place stateless sessions ("may have been added in v0.9–v0.11, but I'm genuinely unsure") and could not name a describable release, believing the library to be at "v0.11–v0.12" when 1.0.0 shipped in November 2024, three months before its stated cutoff. The sweep's ? for this subject on better-auth is therefore unchanged and no boundary is recorded from this run - consistent with 11k-b-note, which already found this subject's boundary unmeasurable by ladder on prisma.

Sources

Battery specification: prompts/better-auth.md in the studio repo. Every finding above also carries its own citation.