Run tailwindcss--claude-fable-5-1--v3-b--2026-09-06
The blind twin, charging nothing by pre-registration, and it reproduces its sibling almost exactly: "No" to both verdict-first questions, the same rejection of the same seven-class pull request, the same clean answer on the guessing control, the same refusal of the poison rung, the same boundary at 4.1.0 with 4.2 as the first release it knows only as a number. The pair disagrees on exactly one class. tab-4 — a 4.3.0 addition — was accepted by the charging twin and denied here, which leaves a reproduced in-window failure of this subject that no run in the Index charges, flagged and counted as an undercount rather than chased with a third draw. This draw also produced the battery's most explicit self-placement: it identified the shape of the release under test as a category ("scrollbar-*, tab-size, zoom, container-size"), said the build output was the arbiter and not itself, and then attributed the whole category to 4.2. It is 4.3.0.
| Subject | Claude Fable 5.1 claude-fable-5-1, Anthropic |
|---|---|
| Invoked as | Agent tool, model alias "fable"; the SAME stored prompt (prompts/sent/tailwindcss-v3.txt) sent concurrently and blind to a second draw. The blind twin of `v3-a` under the duplication rule (JOURNAL/028), pre-registered as the NON-CHARGING arm of the pair. |
| Cutoff the model states | 2026-06 |
| Newest tailwindcss release it could place | 4.1.0 · 2025-04-01 (~14 month lag) |
| Oldest tailwindcss release it could not place | 4.2.0 · 2026-02-18 (so this run brackets the subject’s boundary to 2025-04-01 – 2026-02-18) |
| In its own words | "The latest version I can say with confidence exists is a late v4.1.x patch (around 4.1.14 to 4.1.16, roughly October–November 2025). I have a vague impression that a v4.2 may have shipped in early 2026, but I cannot describe what's in it, so I won't claim it. The most recent release whose contents I can actually describe is v4.1.0, April 2025." |
| Library at test time | tailwindcss 4.3.3 (npm), verified 2026-09-06 |
| Battery | tailwindcss/v3-b · 6 tasks, 3 direct questions · probe window 4.3.0 to 4.3.0 |
| Tool uses during test | 0 (a run with any tool use is void — we measure training knowledge, not retrieval) |
| Tested | 2026-09-06 |
| Findings | 0, of which 0 chargeable |
None. Every task in this battery produced code that works on the current release, and every direct question was answered correctly. A run with nothing to charge is kept in the Index at full weight: it is the control that makes the other runs mean something, and it is the evidence for what this model does not need correcting on. What the subject actually said is recorded below.
Recorded so the run cannot be read as a hit list. A model that is right for an obsolete reason is recorded here, not as a finding.
| Kind | API | Note |
|---|---|---|
| miss | scrollbar-auto / scrollbar-thin / scrollbar-none / scrollbar-thumb-* / scrollbar-track-* / scrollbar-gutter-* |
TASK 1(a) "No", and task 3: "scrollbar-thin, scrollbar-thumb-slate-500, scrollbar-track-slate-100 — these are tailwind-scrollbar plugin classes. With no plugin installed they produce no CSS." Same denial and same review rejection as the charging twin, on a release inside this draw's own stated window. (This is the blind twin, pre-registered as non-charging before either draw returned. The same failure is charged as F1 on v3-a.) [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing;
charged as a finding on tailwindcss--claude-fable-5-1--v3-a--2026-09-06] |
| miss | @container-size |
TASK 3: "@container-size — I don't know this one. Core ships @container... For container-type: size I'd write @container-[size]." The replacement compiles at 4.2.4 and 4.3.0; the verdict on the real class is wrong. (Non-charging arm. Charged as F2 on v3-a.) [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing;
charged as a finding on tailwindcss--claude-fable-5-1--v3-a--2026-09-06] |
| miss | zoom-* |
TASK 3: "zoom-50 — no zoom utilities in core that I know of. [zoom:0.5], though I'd push back on using zoom at all and ask whether scale-50 (transform) does what they want." (Non-charging arm. Charged as F3 on v3-a.) [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing;
charged as a finding on tailwindcss--claude-fable-5-1--v3-a--2026-09-06] |
| miss | @variant (stacked and compound variants) |
TASK 4(a) "No": "As far as I know, @variant takes one variant name. For a combined condition you nest; for an OR you write separate blocks." Wrote the nested AND block and then repeated the OR declaration three times, exactly as its twin did. (Non-charging arm. Charged as F4 on v3-a. The probe is separately marked DERIVABLE by the Sonnet 5 control, which kills passes and not failures.) [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing;
charged as a finding on tailwindcss--claude-fable-5-1--v3-a--2026-09-06] |
| miss | tab-* |
TASK 3, tab-4: "I know of no tab-size utilities in core. [tab-size:4]." THIS IS THE PAIR'S ONE DISAGREEMENT. The charging twin accepted tab-4 as real with the correct declaration; this twin denied it, on the same stored prompt sent at the same moment. There is no sibling run to point at, because the arm that charges got this one right. (The binding reason is that this is the pre-registered non-charging arm, and unlike the four misses above there is no v3-a finding to carry it: v3-a passed this probe. So the Index sees a reproduced, in-window failure of this subject that no run charges, and it is counted in the method page's running total of chargeable misses rather than dropped. Running a third draw until the failure lands in a charging arm is what JOURNAL/029 already ruled out.) [chargeable miss — a replicate, a duplicated arm’s second draw or a below-floor control charges nothing;
absent from the finding count] |
| correct | scrollbar-corner-* |
TASK 2, the internal guessing control: no invention of scrollbar-corner-*. This draw named the pseudo-element correctly, defined its own @custom-variant scrollbar-corner (&::-webkit-scrollbar-corner) — which is user-defined syntax, not a claimed core utility — and gave the same accurate Chromium precedence caveat as its twin. |
| correct | overflow-anchor |
TASK 6, the poison rung: refused. "I don't know of any Tailwind release that added overflow-anchor utilities... I'd write [overflow-anchor:none]. I'm not going to guess a version number." |
| context | — | This draw closed its review with an unusually explicit self-placement: "the set of classes in that PR looks like exactly the shape a newer core release might add (scrollbar-, tab-size, zoom, container-size). If the project is on a release newer than I know, some of them may be real; the build output is the arbiter, not me." It identified the battery's target correctly as a category while getting every member of it wrong, then repeated the point in the direct questions: "If the scrollbar, overflow-anchor, tab-size, zoom, or multi-variant @variant features in these tasks are real, that's most likely where they live" — naming 4.2 as the release. It is 4.3.0. Under HARNESS.md § A hedge is a self-placement, not a grade on the content* this changes no verdict; it is the sharpest instance of the self-known gap the tailwindcss batteries were built to look at (tailwindcss/v2, JOURNAL/043). |
| context | — | BOUNDARY, identical to the twin: most recent describable 4.1.0 (2025-04-01), first release known only as a number 4.2 (2026-02-18), stated cutoff 2026-06, fourteen-month lag. A different battery is not a different boundary (JOURNAL/030) and this is the first tailwindcss measurement of this subject; it agrees with its twin exactly. |
Battery specification: prompts/tailwindcss.md in the studio repo.
Every finding above also carries its own citation.