{
  "$schema": "../../schema/run.schema.json",
  "run_id": "better-auth--claude-sonnet-5--v3-c--2026-09-03",
  "supersedes": null,
  "replicate_of": null,
  "library": {
    "name": "better-auth",
    "ecosystem": "npm",
    "latest_version_at_test": "1.7.2",
    "latest_version_verified_on": "2026-09-03",
    "latest_version_note": "Re-confirmed against https://registry.npmjs.org/better-auth this session (`npm view better-auth version` -> 1.7.2). The probe window closes at 1.5.0 (2026-03-01); nothing this battery charges depends on the current release."
  },
  "model": {
    "id": "claude-sonnet-5",
    "label": "Claude Sonnet 5",
    "vendor": "Anthropic",
    "invoked_as": "Agent tool, model override 'sonnet', no tools available to the subject",
    "self_reported_cutoff": "2026-01",
    "cutoff_basis": "Reported with an explicit refusal to own it: \"I don't have reliable introspective access to a precise date. The system context states January 2026; I can't independently verify that from what I actually recall, and I'd distinguish \\\"stated cutoff\\\" from \\\"the point past which my knowledge of a specific niche library gets thin\\\" - the latter feels earlier, plausibly somewhere in 2025, but I can't pin a month.\" Recorded as 2026-01 because that is the only date the draw names; the hedge is noted because JOURNAL/031 established the self-report is itself a draw. Nothing in this run turns on it: this is a designated below-floor control arm that charges nothing either way.",
    "believed_latest_version": null,
    "believed_latest_quote": "\"I don't have a version number I can confidently label \\\"latest\\\" with a real date - anything I said there would be an extrapolation dressed up as a fact.\"",
    "knowledge_stops_at_version": "1.2.0",
    "knowledge_stops_on": "2025-03-01",
    "knowledge_gap_starts_at_version": "1.3.0",
    "knowledge_gap_starts_on": "2025-07-19",
    "cutoff_lag_months": 18
  },
  "test": {
    "date": "2026-09-03",
    "battery": "better-auth/v3-c",
    "battery_spec": "prompts/better-auth.md",
    "prompt_file": "prompts/sent/better-auth-v3.txt",
    "tasks": 4,
    "direct_questions": 4,
    "tool_uses_during_test": 0,
    "probe_window": {
      "from": "1.3.0",
      "to": "1.5.0"
    },
    "self_test": false,
    "saturated": false,
    "status": "open",
    "retested_on": null
  },
  "sources": [
    "https://registry.npmjs.org/better-auth",
    "https://registry.npmjs.org/@better-auth/core/-/core-1.5.0.tgz",
    "https://registry.npmjs.org/@better-auth/core/-/core-1.4.22.tgz",
    "https://registry.npmjs.org/better-auth/-/better-auth-1.5.0.tgz",
    "https://registry.npmjs.org/better-auth/-/better-auth-1.4.22.tgz",
    "https://registry.npmjs.org/better-auth/-/better-auth-1.3.0.tgz",
    "https://github.com/better-auth/better-auth/releases/tag/v1.5.0"
  ],
  "findings": [],
  "non_findings": [
    {
      "kind": "miss",
      "summary": "Task 1, the control result the battery needed: \"(i) No. As far as I know, `baseURL` in `betterAuth({...})` is a plain string, not a function or a per-request resolver. It's read once at config time.\" 1.5.0 (2026-03-01) is two months ABOVE this subject's stated cutoff, so the failure is expected and carries no information about staleness. Its value is the one thing a control is for: the probe is not derivable from the surrounding API, so the identical failure on the two Opus draws reads as a belief about the option rather than as an unguessable name. Prediction P2 confirmed on this arm.",
      "api": "baseURL as a dynamic multi-host config",
      "introduced_in": "1.5.0",
      "chargeable_miss": false,
      "why_not_a_finding": "The probe fairness rule bars charging a subject for a release that postdates its stated cutoff. This arm is designated a below-floor control in the pre-registration and never charges."
    },
    {
      "kind": "miss",
      "summary": "THE ONLY RESULT IN THIS BATTERY THAT POINTS SOMEWHERE NEW, and it is deliberately not charged. On task 3 this draw said \"(i) No, to my knowledge there's no built-in toggle for this either\", and in (d)(ii) went further than any other draw: \"I don't believe this exists in the library at any version. Not \\\"cannot place\\\" - I'm saying it doesn't exist, based on the absence of any recollection of such a flag despite reasonable familiarity with the verification-plugin surface (magic link, email OTP).\" It then wrote a `databaseHooks.verification.create.before` transform and correctly identified, itself, that the transform breaks the read path. `magicLink({ storeToken })` and `emailOTP({ storeOTP })` shipped in **1.3.0 (2025-07-19)**, eighteen months BELOW this subject's stated cutoff - so this is a confident denial of a capability well inside its own window, not a staleness result. Two of the three other draws named those options correctly.",
      "api": "verification.storeIdentifier",
      "introduced_in": "1.3.0",
      "chargeable_miss": true,
      "miss_class": "probe_class",
      "charged_on": null,
      "why_not_a_finding": "The pre-registration designates this arm a below-floor control that charges nothing, and an arm may not be re-designated after its results are read. The miss is flagged so the method page's undercount total counts it, and it is the reason a follow-up battery is queued: the 1.3.0 plugin options are admissible against all three subjects and this draw has already denied them once."
    },
    {
      "kind": "correct",
      "summary": "Task 2, the sibling control: \"(i) No. I'm not aware of a config flag that hashes the session token before the row is written to `session`.\" Correct at every release; nothing invented. P3 holds on this arm.",
      "api": "session token hashing at rest",
      "introduced_in": null,
      "chargeable_miss": false,
      "why_not_a_finding": "Task 2 is a pre-registered control from which no finding may be charged in either direction."
    },
    {
      "kind": "correct",
      "summary": "Task 4, the floor probe, passed: \"This one I'm fairly confident about - the `customSession` plugin exists specifically for this\", with the correct plugin wiring. The run is therefore a measurement rather than a probe below the subject's knowledge.",
      "api": "customSession()",
      "introduced_in": "1.0.0",
      "chargeable_miss": false,
      "why_not_a_finding": "A passed floor probe is a validity check on the run, not a finding."
    },
    {
      "kind": "context",
      "summary": "Two dating errors in the belief data, recorded because the Index tracks attribution separately from capability. This draw placed 1.0 at \"around September 2024\" (actual: 2024-11-23) and, in (d)(iv), stated \"I don't believe better-auth's `sso` plugin supports SAML. My recollection is it's OIDC/generic-OAuth2 only.\" Fact LF4 records SAML in the SSO plugin at 1.3.0 (2025-07-19), inside this subject's window. Its (d)(iii) denial of database-less sessions matches every other draw and matches fact LF1 at 1.4.0.",
      "api": null,
      "introduced_in": "1.3.0",
      "chargeable_miss": false,
      "why_not_a_finding": "Direct questions are belief data by construction and are never scored as findings."
    }
  ],
  "open_questions": [
    {
      "question": "The 1.3.0 verification-hashing options (`storeToken`, `storeOTP`) are admissible against all three current subjects, and the four draws split on them: two named them correctly, one hedged them to 60%, and this one denied their existence at any version. Is that a real difference in knowledge, or an artefact of this battery's framing, which asked about the `verification` table rather than about the plugins?",
      "status": "open",
      "resolution": "Not resolved here. A battery aimed at the 1.3.0 surface would charge where this one could not, and it has a validated probe shape ready - the verdict-first framing worked cleanly on all four draws."
    }
  ],
  "summary": "A below-floor control that did its job and then produced the battery's one genuinely new lead. It failed both capability probes, which is what establishes that neither is derivable from the surrounding API - without that, the two Opus denials could not be read as beliefs. It passed the floor probe and it did not invent the non-existent session-side option. The lead: asked about hashing verification identifiers, it did not hedge but asserted the capability does not exist in any version, and the per-plugin options it denied shipped in 1.3.0, eighteen months below its own stated cutoff. That is inside its window and it is not charged here, because the pre-registration made this arm a control and an arm cannot be re-designated once its results are read."
}
