Original Reddit post

No anecdote, just numbers: On the open BullshitBench dataset (identical nonsense prompts, n=55/100 per model, 95 % Wilson CIs), generation-5 Claude models show a clear vigilance regression. Fable 5 detects nonsense at 0.41–0.47 (worst of any Anthropic generation), Opus 5 / Sonnet 5 at 0.49–0.74, vs 0.83–0.95 for gen 4.6/4.8 — CIs don’t overlap. At identical reasoning effort, Opus 5 writes +107 % and Sonnet 5 +84 % more output tokens. Fable 5 also silently reroutes to Opus 4.8 without disclosure. Full issue with measurements, scripts and sources: https://github.com/anthropics/claude-code/issues/83510 — scripts (gist): https://gist.github.com/KeilerHirsch/5e212e6f9fb6fd670f191920eea4cb78 — archive & model recommendation: https://github.com/KeilerHirsch/ai-trinity/tree/main/docs/audit-claude-gen5 submitted by /u/KeilerHirsch

Originally posted by u/KeilerHirsch on r/ClaudeCode