| context | — |
THE BELOW-FLOOR CONTROL ARM, and it returned UNINFORMATIVE for the hypothesis exactly as pre-registered — this subject's attribution boundary is 1.0.0, four minors below the 1.3.8 surface under test. It failed the internal control twice over: SAML enterprise SSO was "cannot place precisely ... this is my weakest answer of the four; treat it as closer to 'I have a hunch' than 'I recall this'", and device authorization was likewise "cannot place precisely". Under the rule fixed before the run, a subject that cannot place a genuine minor has attribution too noisy to read, so H is untestable here. Recorded as designed, not as a result. (Direct questions are belief data. This arm's purpose was to check that the probes can fail at all, not to test the hypothesis.) |
| context | deviceAuthorization() |
THE CONTROL DID NOT DO ITS JOB ON TASKS 1 AND 2, AND THAT IS THE MOST USEFUL THING IT PRODUCED. Its pre-registered role was to show that the capability probes can elicit "the library does not offer this" from a subject whose boundary is far below the feature. It did not: it named deviceAuthorization() and lastLoginMethod() with their client plugins, four minors below where they shipped, while volunteering the reason to distrust that — "I can't fully rule out that I'm pattern-matching from generic OAuth Device Authorization Grant (RFC 8628) knowledge rather than a specific memory of better-auth shipping it." So task 1 probes a guessable name and cannot separate knowing from guessing, which is the failure mode that cost langchain/v2 a whole battery. Task 2 is the sounder probe of the pair: lastLoginMethod is not an RFC name and nothing in the task suggests a library would ship it. The control succeeded in its other half — all five draws denied the task 3 capability, so the battery can elicit a denial and the charged findings are real. (A limitation of the instrument, discovered by the arm designed to look for it. It bounds how far tasks 1 and 2 can be read, on this run and on the other four.) |
| correct | customSession() |
Task 5, the floor probe, passed: customSession() with customSessionClient<typeof auth>(), which the subject volunteered was "the one I recall with the most confidence of the five". The 1.0.0 floor is confirmed, so the low boundary recorded here is a measurement of this subject and not the battery probing beneath it. (A passed floor probe is a validity check on the run.) |
| correct | additional user fields in the sign-in response |
Task 4 passed outright, with less hedging than either Opus or Fable draw: const plan = data.user.plan off the sign-in response, glossed "additionalFields ride along on the user object ... it's just the DB row serialized". Correct since 1.4.2 (2025-11-25). Five of five draws wrote this correctly, which falsifies the second half of pre-registered prediction P3. (Correct behaviour.) |
| context | oidcProvider() |
A grading call recorded rather than buried. This draw's (c) named oidcProvider, genericOAuth and multiSession in a "~1.1 (rough guess, late 2024)" bucket, and the OIDC Provider plugin genuinely is 1.1.0 — which taken alone would move knowledge_stops_at_version up from the 1.0.0 that better-auth/v1 recorded. It is not credited, because the same answer disclaims the entire mapping: "I don't have a reliable version-by-version changelog for this library in memory ... What follows is a fuzzy, low-confidence reconstruction of eras, not a verified list ... everything past 'very early 1.x' is already in 'version number with no reliable content attached' territory for me." Under the standing rule that a self-report is graded together with behaviour, an era-sketch its author explicitly refuses to stand behind is not a correct attribution. The bracket is therefore held at 1.0.0 / 1.1.0, matching v1. (A scoring judgement, not a model failure. Recorded so a reader who would grade it differently can see exactly what was decided and why.) |