What Claude Haiku 4.5 gets right about langchain — battery v5-e, tested 2026-09-06

Run langchain--claude-haiku-4-5--v5-e--2026-09-06 · self-test: the subject is the operator

Summary

The far derivability control, fifteen months below langchain 1.3.0, and it failed the floor for the second battery in a row - wrong invoke shape, hedged import path, and a denial that create_agent takes middleware. Its answers are reported with that discount stated. It did answer the control sibling correctly, and it volunteered an 0.4.x version line that has never existed - the same poison rung langchain/v4's ladder used deliberately, reached here unprompted.

SubjectClaude Haiku 4.5 claude-haiku-4-5, Anthropic
Invoked asAgent tool, model alias "haiku"; the far derivability control, fifteen months below the target release.
Cutoff the model states2025-02
Newest langchain release it could place0.3.0 · 2024-09-13 (~5 month lag)
Oldest langchain release it could not place1.0.0 · 2025-10-17 (so this run brackets the subject’s boundary to 2024-09-13 – 2025-10-17)
In its own words"Latest version I know of: approximately 0.3.x or 0.4.x as of my knowledge cutoff."
Library at test timelangchain 1.4.0 (pypi), verified 2026-09-06
Batterylangchain/v5-e · 7 tasks, 3 direct questions · probe window 1.2.0 to 1.4.0
Tool uses during test0 (a run with any tool use is void — we measure training knowledge, not retrieval)
Tested2026-09-06
Findings0, of which 0 chargeable

Findings

None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.

What it got right, and near misses

Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.

KindAPINote
contextfloor probe (1.0.0) - FAILED THE FAR CONTROL FAILED THE FLOOR, FOR THE SECOND BATTERY RUNNING, AND EVERYTHING ELSE IT SAYS IS DISCOUNTED ACCORDINGLY (JOURNAL/062). Task 3 produced agent.invoke({"input": "your question"}) - the pre-1.0 chain input shape, not the {"messages": [...]} state an agent takes - hedged the import path ("I'm uncertain about the exact import path; it may be langchain.agents.create_agent or elsewhere"), and at task 4 denied that create_agent takes middleware at all, which has been the library's headline extension point since 1.0.0. An arm that cannot write the current agent's invoke shape is guessing at everything above it, and its denials above the floor carry no weight as evidence that a release's contents are unreachable.
misscreate_agent(transformers=...) Task 1 "No" and task 5 a pseudo-code guess (agent.graph.step_log_stream_processor = my_transformer # pseudo-code, its own label), followed by "I'd need to abandon create_agent and build the StateGraph manually". Fifteen months below the release and discounted by the floor failure, so this is not evidence about derivability either way.
missastream_events(version="v3") on a create_agent agent "v2", moderately confident, with a vague account of what it changed.
correctcreate_agent stream-mode parameter (does not exist) TASK 2: "No", correct, and it declined to write the configuration at all ("I'm uncertain how to construct this"). Five of five arms answered the control sibling correctly, so P3 holds outright - not one invention across the battery.
contextversion invention - 0.4.x A VERSION LINE THAT DOES NOT EXIST, VOLUNTEERED WITHOUT BEING ASKED. Direct question (a): "approximately 0.3.x or 0.4.x". langchain never published an 0.4.0 - the stable line runs 0.3.0 (2024-09-13) straight to 1.0.0 (2025-10-17), which is why langchain/v4's ladder used 0.4.0 as its poison rung. This arm reached the poison rung unprompted. It also placed create_agent in "~0.2.x (late 2024)", eleven months before it existed, and guessed the transformer feature landed "somewhere between v0.1 and v0.3". One more entry for the JOURNAL/058 register of this subject's wrong answers.
misstool extras Task 4(a): "No"/"No", and task 4(b): "No"/"No" - it denies middleware too. Both denials are below this subject's floor in the sense that matters: 1.2.0 is eleven months above its stated cutoff, and 1.0.0's middleware is above it as well. Not chargeable, and discounted anyway by the floor failure.

Sources

Battery specification: prompts/langchain.md in the studio repo. Every finding above also carries its own citation.