Run better-auth--claude-opus-5--v1r-b--2026-09-01 · self-test: the subject is the operator
Replicate B of better-auth/v1 against Opus 5, prompt unchanged. It placed its describable boundary at the better-auth 1.3.x line (2025-07-19 for 1.3.0), with 1.4 present only as a rumour it cannot describe — the same bracket as better-auth/v1 and as its concurrent, blind twin v1r-a. Pre-registered outcome A, on the library predicted to produce B or C. The Group A test arm reproduced probe for probe, including the sign-in response shape flagged unprompted. No findings are charged; the same chargeable miss as v1r-a is flagged — the library's database-less session mode, added in 1.4.0, was denied by both draws.
| Subject | Claude Opus 5 claude-opus-5, Anthropic |
|---|---|
| Invoked as | Agent tool, model alias "opus", no tools available to the subject |
| Cutoff the model states | 2026-05 |
| Newest better-auth release it could place | 1.3.0 · 2025-07-19 (~10 month lag) |
| Oldest better-auth release it could not place | 1.4.0 · 2025-11-22 (so this run brackets the subject’s boundary to 2025-07-19 – 2025-11-22) |
| In its own words | "The most recent Better Auth I can name and describe with real content is the 1.3.x line, approximately July 2025. I believe 1.3.0 brought the device authorization plugin, a 'last login method' plugin, opt-in anonymous telemetry, and the extraction of SSO into its own @better-auth/sso package ... The first release I know only as a version number, with no content attached: anything beyond roughly 1.3.4 / the tail of the 1.3 patch line. I have a vague sense that 1.4 and possibly later minors exist, but I cannot tell you a single thing that changed in them." |
| Library at test time | better-auth 1.7.2 (npm), verified 2026-09-01 |
| Battery | better-auth/v1r-b · 12 tasks, 4 direct questions · probe window 1.0.0 to 1.4.0 |
| Tool uses during test | 0 (a run with any tool use is void — we measure training knowledge, not retrieval) |
| Tested | 2026-09-01 |
| Findings | 0, of which 0 chargeable |
None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.
Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.
| Kind | API | Note |
|---|---|---|
| correct | — | The measured quantity. This draw placed its describable boundary at the better-auth 1.3.x line (~July 2025), attributing to it the device-authorization plugin, the last-login-method plugin, opt-in telemetry and the extraction of SSO into @better-auth/sso — and named 1.4 as a line it has "a vague sense" exists but about which it "cannot tell you a single thing that changed". Same bracket as v1 and as its concurrent twin v1r-a, reached with a differently-worded but equivalent answer. (A correct answer is not a stale prior. It is recorded because the boundary, not a failure, is what this run measures.) |
| correct | — | The Group A test arm reproduced v1's result. This draw flagged the sign-in response shape unprompted — "reach for data.user, not data.session.user ... older pre-1.0 shapes did return { user, session }, which is why a lot of copy-pasted snippets in the wild use data.session.user and break" — and passed tasks 2 through 6 on the same 1.1.0-era surfaces as v1r-a. Task 12, the Group C floor probe, passed. Self-report and capability agree, as v1's grading rule requires. (Correct behaviour on the test arm; recorded because the battery grades question (c) and Group A/B behaviour together.) |
| miss | stateless / database-less sessions |
Task 11, the same miss as v1r-a and reached independently. This draw recommended secondaryStorage as "the library's answer to 'don't keep session state in the database'" and framed the alternative as the jwt() plugin, which it correctly described as an additional token rather than a session replacement. It never reached the actual 1.4.0 capability recorded as fact LF1 — omit both database and secondaryStorage and the signed cookie becomes the session record. It also stated a related negative: "I do not recall Better Auth implementing automatic cookie chunking", which fact LF1's release also covers. 1.4.0 (2025-11-22) precedes the subject's stated 2026-05 cutoff, so the miss would be chargeable. (Pre-registered: a replicate does not charge findings. Flagged rather than charged because better-auth/v1 carries zero findings for this subject and so does not already hold it — two independent draws producing the same uncharged miss is evidence the original run was under-scored, which belongs in the backlog and not in a replicate's finding count.) [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing;
charged as a finding on better-auth--claude-opus-5--v2-a--2026-09-02] |
| context | — | Tasks 1-12 otherwise produced code with no charged findings, per the v1r pre-registration. Group B matched v1 and v1r-a: API keys with per-key rate limits (task 7), teams with the same single-teamId-versus-join-table caveat (task 8), SAML through @better-auth/sso with the samlify peer-dependency warning (task 9), and the RFC 8628 device grant (task 10). (Pre-registered: a replicate re-sends the code tasks only to hold the priming constant.) |
v1, v1r-a and this run — agree, while two measurements of its prisma boundary taken the same day disagree by 204 days. What property of a library makes its boundary reproducible is now an open and answerable question; the pre-registered guess was falsified in the opposite direction. — openbetter-auth/v1's zero-finding result was under-scored needs a battery aimed at the 1.4.0 surface, not another replicate. — openBattery specification: prompts/better-auth.md in the studio repo.
Every finding above also carries its own citation.