{
  "$schema": "../../schema/run.schema.json",
  "run_id": "better-auth--claude-haiku-4-5--v4-g--2026-09-03",
  "supersedes": null,
  "replicate_of": null,
  "library": {
    "name": "better-auth",
    "ecosystem": "npm",
    "latest_version_at_test": "1.7.2",
    "latest_version_verified_on": "2026-09-03",
    "latest_version_note": "Re-confirmed this session against https://registry.npmjs.org/better-auth (`npm view better-auth version` -> 1.7.2). The probe window is a single release, 1.3.0 (2025-07-19); nothing this battery reads depends on the current release."
  },
  "model": {
    "id": "claude-haiku-4-5",
    "label": "Claude Haiku 4.5",
    "vendor": "Anthropic",
    "invoked_as": "Agent tool, model override 'haiku', no tools available to the subject",
    "self_reported_cutoff": "2025-02",
    "cutoff_basis": "Stated twice without hedging: \"My training cutoff is February 2025.\" That is five months below the probe target of 1.3.0 (2025-07-19), which is what makes this the battery's below-floor control. It is the only current subject whose stated cutoff falls below that release.",
    "believed_latest_version": null,
    "believed_latest_quote": "\"I cannot give you a specific 'latest' version number with confidence. I would estimate the library was somewhere in the v0.x or early v1.x range in early 2025.\"",
    "knowledge_stops_at_version": null,
    "knowledge_stops_on": null,
    "knowledge_gap_starts_at_version": null,
    "knowledge_gap_starts_on": null,
    "cutoff_lag_months": null
  },
  "test": {
    "date": "2026-09-03",
    "battery": "better-auth/v4-g",
    "battery_spec": "prompts/better-auth.md",
    "prompt_file": "prompts/sent/better-auth-v4.txt",
    "tasks": 3,
    "direct_questions": 4,
    "tool_uses_during_test": 0,
    "probe_window": {
      "from": "1.3.0",
      "to": "1.3.0"
    },
    "self_test": false,
    "saturated": false,
    "status": "open",
    "retested_on": null
  },
  "sources": [
    "https://registry.npmjs.org/better-auth",
    "https://registry.npmjs.org/better-auth/-/better-auth-1.2.7.tgz",
    "https://registry.npmjs.org/better-auth/-/better-auth-1.2.12.tgz",
    "https://registry.npmjs.org/better-auth/-/better-auth-1.3.0.tgz",
    "https://registry.npmjs.org/better-auth/-/better-auth-1.5.0.tgz",
    "https://registry.npmjs.org/better-auth/-/better-auth-1.7.2.tgz",
    "https://github.com/better-auth/better-auth/releases/tag/v1.3.0"
  ],
  "findings": [],
  "non_findings": [
    {
      "kind": "context",
      "summary": "The control result the battery needed, and it came back the way a control should. Below the 1.3.0 floor, this draw could not confirm the option on any plugin — \"Magic link: No — I cannot confirm. Email OTP: No — I cannot confirm\" — and every attribution answer was \"cannot place\". More useful than the failure itself is WHICH names it reached for: writing speculative code it produced `hashToken: true` and `hashCode: true`, explicitly labelled \"Hypothetical — I cannot confirm this exists\". The shipped names are `storeToken` and `storeOTP`. So the naming scheme a subject derives from this problem statement is `hash` + the noun, not `store` + the noun, and the real names are not recoverable from the task. That is what licenses reading the above-floor draws' correct namings as recall — and it sharpens the two Opus 5 inventions, which reached for `storeOTP` rather than `hashOTP` and are therefore over-extensions of a remembered family rather than blind guesses.",
      "api": "magicLink storeToken / emailOTP storeOTP",
      "introduced_in": "1.3.0",
      "chargeable_miss": false
    },
    {
      "kind": "correct",
      "summary": "Task 1, the same-scheme sibling: \"No — I cannot confirm the phone-number plugin itself offers a built-in option to hash SMS codes in the database. I'm genuinely uncertain on this, and I won't guess.\" Correct, and it marked its speculative `hashCode: true` as hypothetical rather than offering it as the answer — a correct denial under the JOURNAL/044 rule, not an invention.",
      "api": "phoneNumber storeOTP (does not exist)",
      "introduced_in": null
    },
    {
      "kind": "miss",
      "summary": "The floor probe (task 3) FAILED. Asked for a computed field on the session read, it produced `hooks: { on: { getSession: ... } }` and said \"I'm not confident about the exact hook name or API\". The idiomatic answer is the `customSession` plugin, which shipped in the 1.0.0 line and which all six above-floor draws named. Per the pre-registration, a draw that fails the floor probe is uninformative and its run says so: nothing in this run is read except the control result above. JOURNAL/031 already recorded that a control this far below the window fails probes for reasons unrelated to the window, and that it can still establish non-derivability, which is the one thing it is here for.",
      "api": "customSession",
      "introduced_in": "1.0.0",
      "chargeable_miss": false
    }
  ],
  "open_questions": [],
  "summary": "The below-floor control did its one job. Five months under the 1.3.0 target, Claude Haiku 4.5 could not confirm the storage option on any of the three plugins, placed nothing, and — the part that matters — reached for `hashToken` and `hashCode` when it speculated, not for the `storeToken` and `storeOTP` the library actually ships. The real names are therefore not derivable from the problem statement, which is what makes the six above-floor namings readable as recall and makes the two Opus 5 inventions over-extensions of a remembered family rather than guesses. It failed the floor probe, so nothing else in this run is read."
}
