Claude Fable 5.1 (measured across 1.0.0 to 1.2.0), Claude Fable 5 (measured at 1.0.0), Claude Opus 5 (measured at 1.0.0) and Claude Sonnet 5 (measured across 0.3.0 to 1.0.0) — nothing measured here sits above them, and nothing separates them from each other. Of the 10 pairs among the 5 subjects with a measured langchain boundary, 3 are separated at all (Claude Fable 5.1 above Claude Haiku 4.5, Claude Fable 5 above Claude Haiku 4.5 and Claude Opus 5 above Claude Haiku 4.5); the rest overlap, and an overlap is a tie rather than a rank. Read "best" precisely: it means best at version attribution — the newest langchain release whose contents the subject can correctly attribute to that release — not most capable at writing langchain code. A subject past its boundary often writes working code with a newer API while naming the wrong release for it. The order is by measured boundary and by nothing else — the finding counts below are deliberately not the ranking, because probing effort is not equal: Claude Sonnet 5 has 8 runs published on langchain against Claude Haiku 4.5's 2, and a subject probed more often has more chances to be caught. Claude Fable 5's interval is a single draw rather than a range, which is narrow because it was measured once, not because it is stable.
Ranked by one rule, applied to every page of this kind: a subject is placed above another only where the whole of its measured boundary interval sits above the whole of the other's (its lowest measured boundary above the other's highest); every overlap is a tie, and finding counts are not the order. Every figure below is read from the dataset at build time; nothing on this page is written by hand. langchain 1 means the major that began at 1.0.0 (2025-10-17); langchain was at 1.4.0 when this index last verified against it (2026-09-06).
| Subject | Boundary measured | Runs bracketing it | Stated cutoff | Findings charged |
|---|---|---|---|---|
| Claude Fable 5.1
nothing measured above it |
1.0.0 2025-10-17 1.1.0 x2 2025-11-24 1.2.0 2025-12-15 |
4 of 4 | 2026-06 ~6 to ~8 months behind |
2 2x S3 deprecated |
| Claude Fable 5
nothing measured above it |
1.0.0 2025-10-17 |
1 of 3 | 2026-01 ~3 months behind |
4 3x S2 silently-wrong, 1x S4 wrong-metadata |
| Claude Opus 5
nothing measured above it |
1.0.0 x4 2025-10-17 |
4 of 6 | 2026-05 ~7 months behind |
6 3x S2 silently-wrong, 2x S3 deprecated, 1x S4 wrong-metadata |
| Claude Sonnet 5
nothing measured above it |
0.3.0 x4 2024-09-13 1.0.0 x2 2025-10-17 |
6 of 8 | 2026-01 ~3 to ~16 months behind |
8 1x S1 breaks-build, 4x S2 silently-wrong, 2x S3 deprecated, 1x S4 wrong-metadata |
| Claude Haiku 4.5 | 0.2.0 2024-05-20 0.3.0 2024-09-13 |
2 of 2 | 2025-02 ~5 to ~9 months behind |
0 |
Where a subject shows more than one boundary, those are separate draws of the same measurement, printed rather than averaged: the mean of two draws is a release nothing measured. A subject with one bracketing run has an interval one release wide because it was measured once, not because it is stable.
A subject is placed above another only where the whole of its measured interval sits above the whole of the other’s. Overlaps are ties. That is a partial order, not a table with a winner, and this is what it separates on langchain:
Ties are not merged into tiers. A greedy sweep would put two subjects in the same tier because a third with a wide spread overlaps both, and would print a relation nothing measured.
Finding counts measure what probing found, and probing effort is not equal across subjects. Sorting by them would publish testing effort as model quality, so they are printed and not ranked on.
| Subject | Published runs on langchain | Findings charged | Findings per run |
|---|---|---|---|
| Claude Sonnet 5 | 8 | 8 | 1.0 |
| Claude Opus 5 | 6 | 6 | 1.0 |
| Claude Fable 5.1 | 4 | 2 | 0.5 |
| Claude Fable 5 | 3 | 4 | 1.3 |
| Claude Haiku 4.5 | 2 | 0 | 0.0 |
Findings per run is printed for the same reason: it is the count adjusted for effort, and it is still not the ranking — the batteries differ in what they probe, so two runs are not one unit.
The same data, asked the other way round — one page per subject, with its boundary measurements, its findings on langchain, and how repeatable the boundary was.