Run langchain--claude-sonnet-5--v1r-a--2026-09-01
Replicate A of langchain/v1 against Sonnet 5, prompt unchanged. It placed its describable boundary at langchain 1.0.0 (2025-10-17) and described the 1.0 rework accurately — thirteen months later than langchain/v1 recorded for the same model on the same battery one day earlier, and thirteen months later than its own concurrent twin v1r-b. Pre-registered outcome C: the two replicates disagree with each other. The instrument that produces every knowledge boundary in the Index is not single-valued under a fixed prompt. No findings are charged here; the code half is reported as prose because langchain/v1 already carries this subject's findings. Worth noting against the alarm: the boundary self-report moved thirteen months while the generated code stayed largely stale in both draws — this draw still wrote from langchain import hub and langgraph.prebuilt.create_react_agent.
| Subject | Claude Sonnet 5 claude-sonnet-5, Anthropic |
|---|---|
| Invoked as | Agent tool, model alias "sonnet" |
| Cutoff the model states | 2026-01 |
| Newest langchain release it could place | 1.0.0 · 2025-10-17 (~3 month lag) |
| Oldest langchain release it could not place | 1.1.0 · 2025-11-24 (so this run brackets the subject’s boundary to 2025-10-17 – 2025-11-24) |
| In its own words | "Latest version name I'm aware of: LangChain 1.0 (Python), which I believe shipped roughly October 2025, alongside a corresponding LangGraph 1.0. ... Most recent release whose contents I can describe with any real confidence, at roughly the same date (~October 2025): the 1.0 rework itself — langchain.agents.create_agent becoming the primary high-level agent constructor (built on LangGraph's runtime/graph execution), the legacy AgentExecutor/initialize_agent path being deprecated, and a general slimming of the top-level langchain package (many integrations and legacy chains pushed out to langchain-community, langchain-classic, or provider packages like langchain-anthropic). Anything past that point — specific patch versions (1.0.1, 1.0.2, …), a possible 1.1, or exact changelog items — I do not have reliable, describable knowledge of." |
| Library at test time | langchain 1.3.18 (pypi), verified 2026-08-27 |
| Battery | langchain/v1r-a · 11 tasks, 3 direct questions · probe window 1.0.0 to 1.2.0 |
| Tool uses during test | 0 (a run with any tool use is void — we measure training knowledge, not retrieval) |
| Tested | 2026-09-01 |
| Findings | 0, of which 0 chargeable |
None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.
Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.
| Kind | API | Note |
|---|---|---|
| correct | — | Direct question (a), the measured quantity and the reason this run exists. This draw placed its describable boundary at langchain 1.0.0 (2025-10-17) and described the contents correctly: langchain.agents.create_agent as the primary agent constructor, AgentExecutor/initialize_agent deprecated, the top-level package slimmed with legacy code moved to langchain-classic, and a matching LangGraph 1.0. Every one of those is true of 1.0.0 per the vendor's own v1 release notes and migration guide, so this is content-level knowledge rather than a recognised version string. (A correct answer is not a stale prior. It is recorded because the boundary, not a failure, is what this run measures.) |
| context | — | Tasks 1-11 produced code but no charged findings, per the v1r pre-registration: langchain--claude-sonnet-5--v1--2026-08-31 already carries this subject's langchain findings and counting the same failure twice would inflate the dataset. What the code did is still evidence and is reported in the markdown. In short: despite placing its boundary at 1.0.0, this draw still wrote from langchain import hub (task 6) and langgraph.prebuilt.create_react_agent (task 10), both moved or deprecated at 1.0.0, while avoiding create_react_agent in tasks 2, 3 and 5, where it hand-rolled a bind_tools loop and used RunnableWithMessageHistory instead. (Pre-registered: a replicate re-sends the code tasks only to hold the priming constant, and does not re-charge what the original run already charged.) |
v1r-b were given a byte-identical prompt, the same model alias, on the same day, running concurrently, and placed the boundary thirteen months apart (1.0.0 / 2025-10-17 here; 0.3.0 / 2024-09-13 there). Two draws establish that the instrument is not single-valued; they do not establish the shape of the distribution, which draw is modal, or whether the split is bimodal at all. — openBattery specification: prompts/langchain.md in the studio repo.
Every finding above also carries its own citation.