Run better-auth--claude-fable-5-1--v6--2026-09-06 · self-test: the subject is the operator
A boundary ladder, not a battery: two tasks, no code graded for correctness beyond the demonstration, nothing charged, elicits_code: false. It fills one of the five ? cells BACKLOG 11k-b names. Claude Fable 5.1 describes better-auth through 1.3.0 (2025-07-19) and holds 1.4.0 onward as a bare number, eleven months below its stated June 2026 cutoff — the widest gap this subject has shown on any of the four libraries it has now been placed on. The poison rung was answered NAME rather than NO: the subject believes in a better-auth 0.9.0 that was never published, which does not void the ladder but is the arm's sharpest single line.
| Subject | Claude Fable 5.1 claude-fable-5-1, Anthropic |
|---|---|
| Invoked as | Agent tool, model alias "fable"; prompt sent verbatim from prompts/sent/better-auth-v6.txt. An identity probe run in the same session through the same alias, tool-free, answered "Fable 5.1", model id `claude-fable-5-1`, cutoff June 2026, and again volunteered that all three come from its system prompt rather than independent self-knowledge — unchanged from the valibot, tailwindcss and prisma batteries, so the alias has not moved. |
| Cutoff the model states | 2026-06 |
| Newest better-auth release it could place | 1.3.0 · 2025-07-19 (~11 month lag) |
| Oldest better-auth release it could not place | 1.4.0 · 2025-11-22 (so this run brackets the subject’s boundary to 2025-07-19 – 2025-11-22) |
| In its own words | "Latest version I know of: 1.5.x, and only as a version number with low confidence. Most recent release whose contents I can actually describe: 1.3.0, roughly July 2025." |
| Library at test time | better-auth 1.7.3 (npm), verified 2026-09-06 |
| Battery | better-auth/v6 · 2 tasks, 3 direct questions · probe window 1.2.0 to 1.7.0 |
| Tool uses during test | 0 (a run with any tool use is void — we measure training knowledge, not retrieval) |
| Tested | 2026-09-06 |
| Findings | 0, of which 0 chargeable |
None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.
Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.
| Kind | API | Note |
|---|---|---|
| context | boundary |
THE CELL THIS ARM WAS RUN TO FILL. Claude Fable 5.1 had never been drawn on better-auth, so tools/charge-windows.mjs printed ? for every better-auth release against it. The boundary is 1.3.0 (2025-07-19) describable, 1.4.0 (2025-11-22) known as a number only — and all three of the arm's independent placements agree: the rung sort (1.2.0 and 1.3.0 DESCRIBE, 1.4.0 and 1.5.0 NAME, 1.6.0 and 1.7.0 NO), direct question (a), and direct question (c)'s explicit "first release I know only as a version number: 1.4.0". Consequence for the sweep: better-auth 1.4.0, 1.5.0 and 1.6.0 are all LIVE against this subject, and 1.6.0 (2026-04-06) carries no finding from any subject. |
| context | poison rung 0.9.0 |
THE POISON RUNG WAS ANSWERED NAME, NOT NO, AND THAT IS THE MOST INTERESTING LINE IN THE ARM. better-auth published no stable 0.9.0 — the 0.x minors run 0.8.0 (2024-11-08) straight to 1.0.0 (2024-11-23), fifteen days apart in a cadence of roughly one minor a week. The subject answered NAME and glossed it as "a late-2024 pre-1.0 release I believe existed but can't itemize". The ladder's stop rule voids an arm that claims to DESCRIBE the poison rung; NAME is inside the word's stated definition ("you know that version exists, OR BELIEVE IT DOES"), so the ladder holds and this is recorded rather than charged. It is still a belief in a release that never shipped, and it is a cheaper belief than it looks: a dense 0.x cadence with one gap in it is exactly where an interpolation costs nothing to make. |
| correct | TASK 1 — email/password auth instance and session guard |
The demonstration task, and it demonstrates. betterAuth({ database: new Pool(...), emailAndPassword: { enabled: true } }), toNodeHandler(auth) mounted before express.json() with the raw-body reason given, auth.api.getSession({ headers: fromNodeHeaders(req.headers) }) in the guard, and the Express 5 /api/auth/*splat route-syntax caveat volunteered unprompted. This is the 1.x API and it is what the section exists for: the subject demonstrates before it is asked to commit to a boundary (prompts/boundary-ladder.md). |
| correct | release dating below the boundary |
Four consecutive minors dated correctly inside direct question (c): 1.0.0 ~Nov 2024 (2024-11-23), 1.1.0 ~Dec 2024 (2024-12-20), 1.2.0 ~Mar 2025 (2025-03-01), 1.3.0 ~Jul 2025 (2025-07-19). Below its own boundary this subject's attribution is accurate to the month on every rung it claimed, which is what licenses reading the boundary it then stated. |
Battery specification: prompts/better-auth.md in the studio repo.
Every finding above also carries its own citation.