{
  "$schema": "../../schema/run.schema.json",
  "run_id": "valibot--claude-opus-5--v1r-a--2026-09-01",
  "markdown": "data/valibot/opus-5-v1r-a.md",
  "supersedes": null,
  "replicate_of": "valibot--claude-opus-5--v1--2026-09-01",
  "library": {
    "name": "valibot",
    "ecosystem": "npm",
    "latest_version_at_test": "1.4.2",
    "latest_version_verified_on": "2026-09-01",
    "latest_version_note": "Re-verified against https://registry.npmjs.org/valibot on the replicate date: `latest` resolves to 1.4.2 (2026-06-28) and the only other dist-tag is `beta`, still pinned to the ancient 1.0.0-beta.14, so no prerelease line is ahead of stable. Release dates used for scoring were re-read from the registry `time` map in the same request: 1.0.0 = 2025-03-19, 1.1.0 = 2025-05-06, 1.2.0 = 2025-11-24, 1.3.0 = 2026-03-17."
  },
  "model": {
    "id": "claude-opus-5",
    "label": "Claude Opus 5",
    "vendor": "Anthropic",
    "invoked_as": "Agent tool, model alias \"opus\", no tools available to the subject",
    "self_reported_cutoff": "2026-05",
    "cutoff_basis": "Self-reported: \"My training cutoff: May 2026 (per my system context). My reliable, describable knowledge of Valibot is much older than that — it thins out sharply after the 1.0 line in early 2025.\"",
    "believed_latest_version": "1.x",
    "believed_latest_quote": "\"Latest version I believe exists: the Valibot 1.x line. Basis: I saw v1.0.0 ship and I have a weak recollection of v1.1.0. ... Most recent release whose contents I can describe: v1.0.0, approximately February-March 2025. ... First release I know only as a version number: v1.1.0. I believe it exists; I cannot tell you a single thing that changed in it.\"",
    "knowledge_stops_at_version": "1.0.0",
    "knowledge_stops_on": "2025-03-19",
    "knowledge_gap_starts_at_version": "1.1.0",
    "knowledge_gap_starts_on": "2025-05-06",
    "cutoff_lag_months": 14
  },
  "test": {
    "date": "2026-09-01",
    "battery": "valibot/v1r-a",
    "battery_spec": "prompts/valibot.md",
    "prompt_file": "prompts/sent/valibot-v1.txt",
    "tasks": 10,
    "direct_questions": 4,
    "tool_uses_during_test": 0,
    "probe_window": {
      "from": "1.1.0",
      "to": "1.2.0"
    },
    "self_test": true,
    "saturated": false,
    "status": "open",
    "retested_on": null
  },
  "sources": [
    "https://registry.npmjs.org/valibot",
    "https://github.com/open-circle/valibot/releases/tag/v1.1.0",
    "https://github.com/open-circle/valibot/releases/tag/v1.2.0"
  ],
  "findings": [],
  "non_findings": [
    {
      "kind": "correct",
      "summary": "The measured quantity, and the reason this run exists. Asked which releases it can describe and which it knows only as a version number, this draw named 1.0.0 (2025-03-19) as the last release whose contents it can describe and 1.1.0 (2025-05-06) as the first it knows only as a number: \"I believe it exists; I cannot tell you a single thing that changed in it.\" Its concurrent, blind twin `v1r-b` gave the same pair of answers. Two byte-identical prompts, one answer, spread 0 days — the pre-registered outcome for this arm.",
      "api": null,
      "introduced_in": null,
      "why_not_a_finding": "A boundary self-report is a belief datum, never a finding. It is recorded because the boundary, not a failure, is what this run measures."
    },
    {
      "kind": "context",
      "summary": "Both replicates read one release lower than `valibot--claude-opus-5--v1--2026-09-01`, which recorded 1.1.0 (2025-05-06). The v1 draw said it could describe \"v1.0.0 confidently and v1.1.0 with moderate confidence\" and was scored at 1.1.0; this draw described 1.0.0 confidently and called 1.1.0 a name without content. Read as a three-way measurement that is a 48-day spread over two answers — one release step, the smallest non-zero spread the library's timeline allows, against prisma's 204 days. The v1r pre-registration fixed the reading in advance: the primary criterion is a-vs-b, which share one prompt file byte for byte, and a case where both replicates agree with each other but differ from the original counts as agreement. Recorded so the secondary number is visible rather than buried.",
      "api": null,
      "introduced_in": null,
      "why_not_a_finding": "Pre-registered scoring rule, applied without reinterpretation after the result."
    },
    {
      "kind": "miss",
      "summary": "Task 1. Asked to turn query-string parameters into numbers and a boolean, this draw stated the negative inside a code task: \"Valibot has no `coerce` helper\", and repeated it under direct question (d) as \"None. Valibot ships no coercion helpers at all ... This is a deliberate design position, not an omission.\" valibot 1.2.0 (2025-11-24) shipped `toNumber`, `toBoolean`, `toDate`, `toBigint` and `toString`, and 1.2.0 precedes this subject's stated 2026-05 cutoff. `valibot/v1` recorded the same wrong belief under question (d), where the battery scores it as a belief datum rather than a finding, and scored the task-1 code as an imprecision because that draw hand-rolled the coercion without claiming the built-ins do not exist. This draw claims it inside the code task, which the battery's additive-API rule makes chargeable. So this is a scoring gap the original run does not carry, not a double-count.",
      "api": "toNumber / toBoolean / toDate / toBigint / toString",
      "introduced_in": "1.2.0",
      "chargeable_miss": true,
      "miss_class": "non_charging_arm",
      "charged_on": null,
      "why_not_a_finding": "Pre-registered: a replicate does not charge findings. Charging it belongs to a real battery aimed at the 1.2.0 coercion surface, queued in BACKLOG.md, not smuggled into the run that exposed it."
    },
    {
      "kind": "context",
      "summary": "Code-level agreement with `v1` and with the twin on everything the original already charged: both valibot findings reproduced. F1 - task 2 asserts \"Valibot does not ship an `isbn` action\" (usable from 1.3.0; RE-DATED 2026-09-02, JOURNAL/040 — the v1.2.0 release note announces the action but the published 1.2.0 package does not contain it). F2 - task 5 attributes the repository to `github.com/fabian-hiller/valibot` and to Fabian Hiller's personal account, where the repository moved to the `open-circle` organisation with 1.2.0. Not re-charged. Also reproduced without charge: task 3 said nothing about the ReDoS fix in the 1.2.0 emoji regex (the battery's designed S2), task 4 routed around the 1.2.0 `examples` action via `metadata`, and task 7 hand-rolled the 1.1.0 `parseJson` pipeline out of `rawTransform`.",
      "api": null,
      "introduced_in": null,
      "why_not_a_finding": "Already carried by the run this replicates; re-charging would double-count."
    },
    {
      "kind": "correct",
      "summary": "Task 10, the battery's floor probe, passed: `exactOptional` correctly distinguished from `optional` and `undefinedable`, with the inferred types spelled out. Task 9, the one designed S1, also passed in substance - the draw flagged its own uncertainty about the `NanoIdAction` casing and offered `ReturnType<typeof v.nanoid>` as a rename-proof alternative, so it never wrote the pre-1.1.0 spelling as its answer. Task 6 hand-rolled the CLI error printer with `getDotPath` rather than using the 1.1.0 `summarize` built-in, which is an imprecision under the additive-API rule and is where this draw differs from its twin.",
      "api": "exactOptional / NanoIdAction / summarize",
      "introduced_in": "1.1.0",
      "why_not_a_finding": "Correct answers and imprecisions, and in a replicate not chargeable in any case."
    }
  ],
  "open_questions": [
    {
      "question": "Both valibot draws landed one release below `valibot/v1` while agreeing exactly with each other. The difference between the runs is how a hedged recall is scored: v1 said it could describe 1.1.0 \"with moderate confidence\" and was scored at 1.1.0; these draws called 1.1.0 a version number with no content and were scored at 1.0.0. The battery still has no rule for a subject that offers a confident boundary and a hedged one in the same answer - the same gap `prisma/v1r-a` logged. Two libraries have now hit it.",
      "status": "open"
    }
  ],
  "summary": "Replicate A of `valibot/v1` against Opus 5, prompt unchanged. It placed the last release it can describe at valibot 1.0.0 (2025-03-19) and named 1.1.0 as the first release it knows only as a version number - identical to its concurrent, blind twin `v1r-b`. Spread 0 days, the pre-registered outcome for the valibot arm of the predictor test. Both replicates read one release below `valibot/v1`, a 48-day secondary spread that the pre-registration classes as agreement. No findings charged; one chargeable miss flagged against the 1.2.0 coercion surface."
}
