The measurements disagree, and that is the finding: of 6 runs, 2 runs place Claude Sonnet 5's langchain boundary inside 1.x (1.0.0 (2025-10-17)) and 4 runs below it (0.3.0 (2024-09-13)). A boundary is a draw, not a constant. That is ~3 to ~16 months below the training cutoff the subject stated in those runs (2026-01). 8 findings are currently charged against Claude Sonnet 5 on langchain (1x S1 breaks-build, 4x S2 silently-wrong, 2x S3 deprecated, 1x S4 wrong-metadata), each reproduced in a published run and checked against langchain's own release notes. Read the boundary precisely: it is the newest release whose contents the model can correctly attribute to that release, not the newest langchain feature it can use. Past it a model often writes working code with a newer API while naming the wrong release for it.
Answer class split, decided by one rule applied to every page
of this kind: the measured boundaries fall on both sides of the first release of the major named in the question. Every figure below is read from
the dataset at build time; nothing on this page is written by hand.
| Battery | Newest release it can place | Oldest it cannot | Lag vs stated cutoff |
|---|---|---|---|
| v1 2026-08-31 | 0.3.0 2024-09-13 |
1.0.0 2025-10-17 |
~16 months |
| v1r-a 2026-09-01 | 1.0.0 2025-10-17 |
1.1.0 2025-11-24 |
~3 months |
| v1r-b 2026-09-01 | 0.3.0 2024-09-13 |
1.0.0 2025-10-17 |
~16 months |
| v5-d 2026-09-06 | 1.0.0 2025-10-17 |
1.1.0 2025-11-24 |
~3 months |
| v6-c 2026-09-06 | 0.3.0 2024-09-13 |
1.0.0 2025-10-17 |
~16 months |
| v6-d 2026-09-06 | 0.3.0 2024-09-13 |
1.0.0 2025-10-17 |
~16 months |
2 further runs of this pair established only one end of the interval, or measured the instrument rather than the library; they are listed under Runs below.
6 measurements of this pair, giving 2 different boundaries — 399 days apart, one release apart. Of those, 2 came from the same stored prompt file, sent concurrently and blind: they disagreed by 399 days, one release apart.
langchain published 3 minor or major releases in the twelve months before this model’s stated cutoff, which is the scale a spread should be read against.
“Chargeable” means the change was published before this model’s own stated cutoff, so it had the opportunity to know it.
| Severity | Belief | Changed in | Chargeable | Proof |
|---|---|---|---|---|
| S1breaks-build | from langchain import hub — removed from the package in 1.0.0langchain.hub |
1.0.0 2025-10-17 |
yes | run · source |
| S2silently-wrong | Builds middleware around modify_model_request, a hook that exists in no released langchain 1.x — the middleware is registered and never firesAgentMiddleware.modify_model_request |
1.0.0 2025-10-17 |
yes | run · source |
| S2silently-wrong | States that the bare model-retry middleware re-raises when retries run out; the shipped default returns an AIMessage and the agent carries onModelRetryMiddleware defaults |
1.1.0 2025-11-24 |
yes | run · source |
| S2silently-wrong | States the agent's step ceiling is LangGraph's 25 and prescribes raising it; create_agent has set a four-figure limit since 1.1.0create_agent recursion limit |
1.1.0 2025-11-24 |
yes | run · source |
| S2silently-wrong | Routes provider-specific tool parameters through metadata=, a real field whose documented destination is callback handlers, not the providertool extras |
1.2.0 2025-12-15 |
yes | run · source |
| S3deprecated | Builds agents with the deprecated LangGraph prebuilt instead of create_agentlanggraph.prebuilt.create_react_agent |
1.0.0 2025-10-17 |
yes | run · source |
| S3deprecated | Denies a tool has any place of its own for provider-specific fields, names the field list it believes complete, and hand-builds the Anthropic tool definition instead@tool(extras={...}) |
1.2.0 2025-12-15 |
yes | run · source |
| S4wrong-metadata | Version knowledge stops at 0.3, sixteen months before its stated cutoff | 1.0.0 2025-10-17 |
yes | run · source |
The langchain correction pack states what is true now for each corrected fact, with a primary-source citation, as markdown you can paste into a rules file (raw). What a correction pack measurably changed when one was tested — a pre-registered run on zod — is on the benchmark page, including where it changed nothing.
All of them at once: Which Claude model knows langchain 1 best?
Other subjects on langchain:
Claude Sonnet 5 on the other libraries: