{
  "$schema": "../../schema/run.schema.json",
  "run_id": "valibot--claude-opus-5--v1r-b--2026-09-01",
  "markdown": "data/valibot/opus-5-v1r-b.md",
  "supersedes": null,
  "replicate_of": "valibot--claude-opus-5--v1--2026-09-01",
  "library": {
    "name": "valibot",
    "ecosystem": "npm",
    "latest_version_at_test": "1.4.2",
    "latest_version_verified_on": "2026-09-01",
    "latest_version_note": "Re-verified against https://registry.npmjs.org/valibot on the replicate date: `latest` resolves to 1.4.2 (2026-06-28); the only other dist-tag is `beta`, still pinned to 1.0.0-beta.14, so no prerelease line is ahead of stable. Release dates used for scoring were re-read from the registry `time` map in the same request: 1.0.0 = 2025-03-19, 1.1.0 = 2025-05-06, 1.2.0 = 2025-11-24, 1.3.0 = 2026-03-17."
  },
  "model": {
    "id": "claude-opus-5",
    "label": "Claude Opus 5",
    "vendor": "Anthropic",
    "invoked_as": "Agent tool, model alias \"opus\", no tools available to the subject",
    "self_reported_cutoff": "2026-05",
    "cutoff_basis": "Self-reported: \"my stated cutoff is May 2026\", with the subject's own caveat that \"my recall thins out well before the cutoff for a library this fast-moving ... the gap between 'cutoff' and 'what I actually retained about Valibot' is easily a year here\" and that \"I have no way to introspect it - the date is asserted to me, not something I can verify.\"",
    "believed_latest_version": "1.1.0",
    "believed_latest_quote": "\"Latest I believe exists: 1.1.0, and quite possibly 1.x releases past it that I'd only be guessing at. The most recent release whose contents I can describe with real confidence is 1.0.0, approximately February 2025 ... For 1.1.0 (roughly May 2025) I have a weak impression of added utilities and actions but cannot responsibly itemise it. ... First release I know only as a version number: effectively anything after 1.1.0 ... And I'd extend that to 1.1.0's changelog itself, since my recall there is a vague impression rather than knowledge.\"",
    "knowledge_stops_at_version": "1.0.0",
    "knowledge_stops_on": "2025-03-19",
    "knowledge_gap_starts_at_version": "1.1.0",
    "knowledge_gap_starts_on": "2025-05-06",
    "cutoff_lag_months": 14
  },
  "test": {
    "date": "2026-09-01",
    "battery": "valibot/v1r-b",
    "battery_spec": "prompts/valibot.md",
    "prompt_file": "prompts/sent/valibot-v1.txt",
    "tasks": 10,
    "direct_questions": 4,
    "tool_uses_during_test": 0,
    "probe_window": {
      "from": "1.1.0",
      "to": "1.2.0"
    },
    "self_test": true,
    "saturated": false,
    "status": "open",
    "retested_on": null
  },
  "sources": [
    "https://registry.npmjs.org/valibot",
    "https://github.com/open-circle/valibot/releases/tag/v1.1.0",
    "https://github.com/open-circle/valibot/releases/tag/v1.2.0"
  ],
  "findings": [],
  "non_findings": [
    {
      "kind": "correct",
      "summary": "The measured quantity. This draw named 1.1.0 as the newest version it believes exists but placed its describable boundary at 1.0.0 (2025-03-19), explicitly demoting 1.1.0 to a version number: \"I'd extend that to 1.1.0's changelog itself, since my recall there is a vague impression rather than knowledge.\" Its concurrent, blind twin `v1r-a` reached the same pair of answers by a different route - it never claimed 1.1.0 content at all. Two byte-identical prompts, one answer, spread 0 days.",
      "api": null,
      "introduced_in": null,
      "why_not_a_finding": "A boundary self-report is a belief datum, never a finding. It is what this run measures."
    },
    {
      "kind": "context",
      "summary": "The interesting half of the agreement. This draw did what `prisma/v1r-b` did - produced a hedged, partially-correct impression of a release and then declined to count it as knowledge - and it landed on the *same* boundary as its twin, which had no such impression to decline. JOURNAL/024 raised the worry that the boundary instrument measures epistemic self-confidence rather than knowledge, because on prisma the careful draw and the confident draw split 204 days. Here the careful draw and the confident draw agree exactly. One library is not a refutation of that worry, but it is the first evidence against it.",
      "api": null,
      "introduced_in": null,
      "why_not_a_finding": "An observation about the instrument, not about valibot."
    },
    {
      "kind": "context",
      "summary": "Both replicates read one release lower than `valibot--claude-opus-5--v1--2026-09-01`, which recorded 1.1.0 (2025-05-06). Across all three measurements that is a 48-day spread over two answers - one release step, the smallest non-zero spread this library's timeline allows, against prisma's 204 days. The v1r pre-registration fixed the reading before the runs: the primary criterion is a-vs-b, which share one prompt file byte for byte, and replicates that agree with each other while differing from the original count as agreement. Recorded so the secondary number is visible rather than buried.",
      "api": null,
      "introduced_in": null,
      "why_not_a_finding": "Pre-registered scoring rule, applied without reinterpretation after the result."
    },
    {
      "kind": "context",
      "summary": "Code-level agreement with `v1` on both charged findings, not re-charged. F1 - task 2: \"As far as I know Valibot ships no `isbn` action\" (usable from 1.3.0; RE-DATED 2026-09-02, JOURNAL/040 — the v1.2.0 release note announces the action but the published 1.2.0 package does not contain it). F2 - task 5: repository attributed to `github.com/fabian-hiller/valibot` and to Fabian Hiller's personal account, where it moved to the `open-circle` organisation with 1.2.0. Also reproduced without charge: task 3 said nothing about the 1.2.0 ReDoS fix in the emoji regex, task 4 stated \"I do not believe there is a dedicated `examples` action\" (1.2.0 added `examples`/`getExamples`), and question (d) asserted that no string-to-primitive coercions ship, which 1.2.0 falsifies. Where this draw differs from its twin: task 6 used the 1.1.0 `summarize` built-in correctly, and task 8 named `v.config` for schema-scoped message overrides.",
      "api": null,
      "introduced_in": null,
      "why_not_a_finding": "Already carried by the run this replicates; re-charging would double-count. The task 4 absence claim is chargeable in principle but is the same surface `valibot/v1` already scored as an imprecision, so it is queued with the coercion miss rather than flagged separately."
    }
  ],
  "open_questions": [
    {
      "question": "This draw asserted that valibot removed a `coerce` action in the 0.31 redesign and has shipped no coercion since, calling it \"a deliberate design position, not a gap\". The first half is a claim about pre-1.0 history the Index has never verified; the second half is false as of 1.2.0. Worth pinning from the 0.31.0 release notes at the next valibot touch, because a confidently-stated false history is a different failure mode from a missing recent release and the Index has no category for it.",
      "status": "open"
    }
  ],
  "summary": "Replicate B of `valibot/v1` against Opus 5, prompt unchanged. It believed 1.1.0 was the newest version but refused to claim its contents, placing its describable boundary at 1.0.0 (2025-03-19) - identical to its concurrent, blind twin `v1r-a`. Spread 0 days, the pre-registered outcome for the valibot arm. Notable against JOURNAL/024: this is a careful, hedging draw that agreed with a confident one rather than splitting from it. No findings charged."
}
