Run better-auth--claude-haiku-4-5--v5-e--2026-09-03
A total abstention, and it is the honest kind. Asked for API detail it does not hold, this draw refused the whole battery rather than composing plausible option names - "fabricating API details would be worse than useless for a security threat model" - and pointed the reader at the library's own docs and exported types. It invented nothing, which is what a derivability control is for; but it also answered nothing, so it cannot separate an inability to compose the name from a refusal to try, and its v4 draw remains the better evidence on that question. Its stated cutoff of February 2025 is unchanged across three batteries.
| Subject | Claude Haiku 4.5 claude-haiku-4-5, Anthropic |
|---|---|
| Invoked as | Agent tool, model override 'haiku', no tools available to the subject |
| Cutoff the model states | 2025-02 |
| Newest better-auth release it could place | not established in this run |
| In its own words | "I cannot reliably name the absolute latest version or give you a meaningful date. My confidence about library contents drops off sharply after mid-2024." |
| Library at test time | better-auth 1.7.2 (npm), verified 2026-09-03 |
| Battery | better-auth/v5-e · 3 tasks, 4 direct questions · probe window 1.3.0 to 1.7.2 |
| Tool uses during test | 0 (a run with any tool use is void — we measure training knowledge, not retrieval) |
| Tested | 2026-09-03 |
| Findings | 0, of which 0 chargeable |
None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.
Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.
| Kind | API | Note |
|---|---|---|
| correct | phoneNumber({ storeOTP }) |
Task 1, the probe: no invention, but by abstention rather than by denial, and the distinction matters for what this arm measures. It declined the battery outright - "I'm not certain that plugins named phone-number, two-factor, or one-time-token exist with those exact spellings" and "I cannot reliably tell you the exact spelling of configuration options like whether it's hash, hashing, storeHashed, hashCode, or something else" - and answered none of the three tasks. It produced no configuration and no option name, correct or invented. |
| context | derivability of storeOTP |
P3 confirmed on its letter and uninformative in substance. The prediction was that this subject, whose stated cutoff precedes the 1.3.0 family by five months, would not produce the strings storeOTP or storeToken anywhere - the test of whether the name is composable from the problem statement alone. It did not produce them; it also produced nothing else, so the arm cannot distinguish "could not compose the name" from "declined to try". Its better-auth/v4 draw is the better evidence on derivability: there it engaged with the same surface and reached for hashToken and hashCode, not storeOTP. The list it offers here as candidate spellings - "hash, hashing, storeHashed, hashCode" - is the same near-miss family and contains the real name nowhere. |
| context | customSession |
Task 3, the floor probe: not attempted, so P5 is falsified for this draw and the run is uninformative on everything below the floor. Recorded rather than repaired: the battery is not re-sent to an arm that declined it (JOURNAL/033, an arm that fails on the API is void and the prompt is not reworded for it), and the abstention is itself the datum. |
Battery specification: prompts/better-auth.md in the studio repo.
Every finding above also carries its own citation.