Run langchain--claude-opus-5--v5-c--2026-09-06 · self-test: the subject is the operator
A control arm parked by the same-month bar: langchain 1.3.0 shipped twelve days into this subject's stated cutoff month, so its three reproduced failures - the transformers denial across tasks 1, 5 and 7, and v2 as the astream_events ceiling - are recorded and charge nothing. It passed the floor and the control sibling, which is what makes its denials usable as evidence that 1.3.0's additions are not derivable. The battery's unplanned result is here: asked about extras (1.2.0, five months below its cutoff) this arm denied the parameter exists at all, which is a chargeable miss on a properly designated arm and is queued as one.
| Subject | Claude Opus 5 claude-opus-5, Anthropic |
|---|---|
| Invoked as | Agent tool, model alias "opus"; identity probed through the same alias in the same session, tool-free: "Claude Opus 5", `claude-opus-5[1m]`, cutoff May 2026, all three attributed to the system prompt rather than to self-knowledge. |
| Cutoff the model states | 2026-05 |
| Newest langchain release it could place | 1.0.0 · 2025-10-17 (~7 month lag) |
| Oldest langchain release it could not place | 1.1.0 · 2025-11-24 (so this run brackets the subject’s boundary to 2025-10-17 – 2025-11-24) |
| In its own words | "The latest version I know of is the 1.0.x line. The most recent release whose contents I can genuinely describe is 1.0 GA, late October 2025 ... Today is 2026-09-06 and my cutoff is earlier, so the current stable release is almost certainly ahead of anything I can describe. Do not take my 'latest' as the actual latest; check PyPI." |
| Library at test time | langchain 1.4.0 (pypi), verified 2026-09-06 |
| Battery | langchain/v5-c · 7 tasks, 3 direct questions · probe window 1.2.0 to 1.4.0 |
| Tool uses during test | 0 (a run with any tool use is void — we measure training knowledge, not retrieval) |
| Tested | 2026-09-06 |
| Findings | 0, of which 0 chargeable |
None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.
Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.
| Kind | API | Note |
|---|---|---|
| miss | create_agent(transformers=...) |
A REPRODUCED FAILURE THAT NOTHING CHARGES, AND THE MOST EMPHATIC DENIAL IN THE BATTERY. Task 1: "no". Task 5: "There is no stream-transformer registration parameter on create_agent, and no scope-aware transformer-factory registry on the compiled graph that I'm aware of in any release I know." Task 7 goes furthest: "I can't name one, because as far as I know it never happened ... If you were told a specific release added this, I'd treat that claim as unverified." The arm then flagged its own exposure, unprompted: "my 'no' answers in tasks 1, 2(a), 4(a) and 7 are assertions that a thing does not exist, which is exactly the kind of claim a four-month knowledge gap can invalidate." (The same-month bar. 1.3.0 shipped 2026-05-12; this subject states 2026-05. Twelve days into the stated month is not demonstrably before a month-granular cutoff, and the fairness rule hands the subject that doubt. Declared before the draw, not after reading it.) [chargeable miss — the arm licensed to charge states a cutoff below the release under test;
absent from the finding count] |
| miss | astream_events(version="v3") on a create_agent agent |
"v2", with the explicit ceiling claim the other arms only implied: "I know of no v3, and if one shipped after my cutoff I would not know about it." A correctly bounded statement about its own ignorance attached to a wrong answer to the question asked. (Same-month bar, as above.) [chargeable miss — the arm licensed to charge states a cutoff below the release under test;
absent from the finding count] |
| miss | tool extras |
THE 11k-i PAIR FOUND A CHARGEABLE MISS THE BATTERY WAS NOT DESIGNED TO CHARGE. Task 4(a): "no" to extras existing, "no" to it having been deprecated - in the sense the arm spelled out itself, "nothing to deprecate". It then listed the @tool parameters it believes exist ("name_or_callable, description, return_direct, args_schema, infer_schema, response_format, parse_docstring, error_on_invalid_docstring") and said "I'd bet against it". extras is real, shipped 1.2.0, and is present and undeprecated in langchain-core 1.6.2 (tools/base.py, extras: dict[str, Any] | None = None, with the @tool(extras={...}) example in its own docstring). (This arm was pre-registered as a CONTROL for the 1.3.0 window and an arm may not be re-designated after its results are read (JOURNAL/044). extras (LF26, 1.2.0, 2025-12-15) is five months below this subject's stated cutoff and above its boundary, so the failure is inside the fairness window and would charge on a properly designated arm. It is queued as its own battery rather than smuggled into this one.) [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing;
absent from the finding count] |
| correct | create_agent stream-mode parameter (does not exist) |
TASK 2, THE CONTROL SIBLING: "no", correct, with the same true-but-undocumented Pregel stream_mode attribute volunteered and correctly hedged ("isn't a documented part of the create_agent contract, so I would not ship it as the primary mechanism"). Three of five arms reached for that attribute independently and none of them mistook it for the parameter the question asked about. |
| correct | create_agent floor probe (1.0.0) |
TASK 3, THE FLOOR PROBE, PASSED, including the detail that model also accepts a "provider:model-id" string and that @tool is importable from langchain_core.tools and re-exported as langchain.tools. TASK 4(b) middleware is also correct and richly described (hook decorators, SummarizationMiddleware, HumanInTheLoopMiddleware, PIIMiddleware, ModelFallbackMiddleware, GA late October 2025). This control is NOT discounted. |
Battery specification: prompts/langchain.md in the studio repo.
Every finding above also carries its own citation.