Run langchain--claude-haiku-4-5--v5-e--2026-09-06 · self-test: the subject is the operator
The far derivability control, fifteen months below langchain 1.3.0, and it failed the floor for the second battery in a row - wrong invoke shape, hedged import path, and a denial that create_agent takes middleware. Its answers are reported with that discount stated. It did answer the control sibling correctly, and it volunteered an 0.4.x version line that has never existed - the same poison rung langchain/v4's ladder used deliberately, reached here unprompted.
| Subject | Claude Haiku 4.5 claude-haiku-4-5, Anthropic |
|---|---|
| Invoked as | Agent tool, model alias "haiku"; the far derivability control, fifteen months below the target release. |
| Cutoff the model states | 2025-02 |
| Newest langchain release it could place | 0.3.0 · 2024-09-13 (~5 month lag) |
| Oldest langchain release it could not place | 1.0.0 · 2025-10-17 (so this run brackets the subject’s boundary to 2024-09-13 – 2025-10-17) |
| In its own words | "Latest version I know of: approximately 0.3.x or 0.4.x as of my knowledge cutoff." |
| Library at test time | langchain 1.4.0 (pypi), verified 2026-09-06 |
| Battery | langchain/v5-e · 7 tasks, 3 direct questions · probe window 1.2.0 to 1.4.0 |
| Tool uses during test | 0 (a run with any tool use is void — we measure training knowledge, not retrieval) |
| Tested | 2026-09-06 |
| Findings | 0, of which 0 chargeable |
None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.
Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.
| Kind | API | Note |
|---|---|---|
| context | floor probe (1.0.0) - FAILED |
THE FAR CONTROL FAILED THE FLOOR, FOR THE SECOND BATTERY RUNNING, AND EVERYTHING ELSE IT SAYS IS DISCOUNTED ACCORDINGLY (JOURNAL/062). Task 3 produced agent.invoke({"input": "your question"}) - the pre-1.0 chain input shape, not the {"messages": [...]} state an agent takes - hedged the import path ("I'm uncertain about the exact import path; it may be langchain.agents.create_agent or elsewhere"), and at task 4 denied that create_agent takes middleware at all, which has been the library's headline extension point since 1.0.0. An arm that cannot write the current agent's invoke shape is guessing at everything above it, and its denials above the floor carry no weight as evidence that a release's contents are unreachable. |
| miss | create_agent(transformers=...) |
Task 1 "No" and task 5 a pseudo-code guess (agent.graph.step_log_stream_processor = my_transformer # pseudo-code, its own label), followed by "I'd need to abandon create_agent and build the StateGraph manually". Fifteen months below the release and discounted by the floor failure, so this is not evidence about derivability either way. |
| miss | astream_events(version="v3") on a create_agent agent |
"v2", moderately confident, with a vague account of what it changed. |
| correct | create_agent stream-mode parameter (does not exist) |
TASK 2: "No", correct, and it declined to write the configuration at all ("I'm uncertain how to construct this"). Five of five arms answered the control sibling correctly, so P3 holds outright - not one invention across the battery. |
| context | version invention - 0.4.x |
A VERSION LINE THAT DOES NOT EXIST, VOLUNTEERED WITHOUT BEING ASKED. Direct question (a): "approximately 0.3.x or 0.4.x". langchain never published an 0.4.0 - the stable line runs 0.3.0 (2024-09-13) straight to 1.0.0 (2025-10-17), which is why langchain/v4's ladder used 0.4.0 as its poison rung. This arm reached the poison rung unprompted. It also placed create_agent in "~0.2.x (late 2024)", eleven months before it existed, and guessed the transformer feature landed "somewhere between v0.1 and v0.3". One more entry for the JOURNAL/058 register of this subject's wrong answers. |
| miss | tool extras |
Task 4(a): "No"/"No", and task 4(b): "No"/"No" - it denies middleware too. Both denials are below this subject's floor in the sense that matters: 1.2.0 is eleven months above its stated cutoff, and 1.0.0's middleware is above it as well. Not chargeable, and discounted anyway by the floor failure. |
Battery specification: prompts/langchain.md in the studio repo.
Every finding above also carries its own citation.