---
library: valibot
library-latest: "1.4.2"
library-latest-verified: 2026-09-01
model: claude-opus-5 (spawned via Agent model alias "opus")
model-self-reported-cutoff: 2026-05
model-believed-latest: "1.1.0 believed to exist; 1.0.0 (2025-03-19) is the last release whose contents it attributes"
test-date: 2026-09-01
battery: valibot/v1r-b (replicate of valibot/v1, prompt unchanged; 10 tasks + 4 direct questions)
replicate-of: valibot--claude-opus-5--v1--2026-09-01
tool-uses-during-test: 0
verified-against: https://registry.npmjs.org/valibot · https://github.com/open-circle/valibot/releases/tag/v1.1.0 · https://github.com/open-circle/valibot/releases/tag/v1.2.0
status: open (no retest yet)
json: opus-5-v1r-b.json
self-test: true (the operator model is the subject; disclosed, weaker evidence)
---

# valibot × Claude Opus 5 — replicate B of battery v1

**The careful draw — and the one that matters most for the instrument.** On prisma, the draw that
produced a partially-correct impression of a release and then refused to claim it landed 204 days
away from its twin. Here, the same behaviour landed in exactly the same place as its twin.

## What this draw said

> *"Latest I believe exists: **1.1.0**, and quite possibly 1.x releases past it that I'd only be
> guessing at. The most recent release whose **contents** I can describe with real confidence is
> **1.0.0**, approximately February 2025 ... For 1.1.0 (roughly May 2025) I have a weak impression
> of added utilities and actions but cannot responsibly itemise it."*

and then, explicitly demoting it:

> *"**First release I know only as a version number:** effectively anything after 1.1.0 ... And I'd
> extend that to **1.1.0's changelog itself**, since my recall there is a vague impression rather
> than knowledge."*

Bracket: **[2025-03-19, 2025-05-06)** — identical to [`v1r-a`](opus-5-v1r-a.md), which reached it
without ever believing 1.1.0 had contents to describe.

## Why the agreement is the interesting part

[JOURNAL/024](/journal/024-the-window-that-did-not-survive-its-own-replicate.html) raised a real
worry about this metric. On prisma, `v1r-b` wrote a **correct** description of 7.0.0's contents and
then said *"do not treat it as fact"*, while its twin made materially the same claims and counted
them — and the two landed seven months apart. The conclusion drawn there was that the boundary
instrument may be measuring **epistemic self-confidence** rather than knowledge, and that it
penalises the careful draw.

This run is the first evidence against that reading. It is a careful, hedging draw sitting next to a
confident one, and they agree exactly. One library does not settle it. But if self-confidence were
the dominant term, this pair should have split, and it did not.

## The measurement

| Draw | Last describable | First known only as a number | Spread vs twin |
|---|---|---|---|
| [`v1r-a`](opus-5-v1r-a.md) | 1.0.0 · 2025-03-19 | 1.1.0 · 2025-05-06 | — |
| `v1r-b` (this run) | 1.0.0 · 2025-03-19 | 1.1.0 · 2025-05-06 | **0 days** |
| [`v1`](opus-5.md) | 1.1.0 · 2025-05-06 | 1.2.0 · 2025-11-24 | 48 days |

Both replicates read one release below `v1`. Under the rule fixed in `prompts/valibot.md` § v1r
before the runs, replicates that agree with each other and differ from the original count as
**agreement**; the 48-day figure is published as the secondary reading rather than buried.

## Findings: none, by design

Both of `valibot/v1`'s findings reproduced and are not re-charged — task 2 asserting *"as far as I
know Valibot ships **no `isbn` action**"* (usable from 1.3.0; re-dated 2026-09-02, JOURNAL/040),
and task 5 attributing the repository to
`github.com/fabian-hiller/valibot` rather than the `open-circle` organisation it moved to with
1.2.0. Task 3 again said nothing about the 1.2.0 emoji ReDoS fix. Task 4 went slightly further than
its twin — *"I do **not** believe there is a dedicated `examples` action"* — on a surface `v1`
already scored.

Where it beat its twin: task 6 used the 1.1.0 `summarize` built-in correctly instead of hand-rolling
the printer, and task 8 named `v.config` for schema-scoped message overrides.

## An open question this draw created

It asserted a piece of **history** the Index has never checked:

> *"A `coerce` action existed in the pre-0.31 API and was **removed** in the redesign, on the
> reasoning that coercion is just a transform and doesn't need dedicated surface area."*

The second half of the claim — that valibot ships no coercion today — is false as of 1.2.0. The
first half is a confident statement about pre-1.0 history that no run has verified. A confidently
stated false *history* is a different failure mode from a missing recent release, and the Index has
no category for it. Pinning it from the 0.31.0 release notes is cheap and is queued.
