Original Reddit post

Not a rant — measured and source-verified. Two open issues consolidate the evidence: Quality regression (Gen 5 / VCST): BullshitBench nonsense detection Gen 5 0.41–0.74 vs Gen 4.6/4.8 0.83–0.95 (95 % Wilson CIs, no overlap), +84–107 % verbosity at identical reasoning effort, silent rerouting (Fable 5 → Opus 4.8). Model pinning is broken by design: 4 bypass vectors — /model internal write (0 hook events), undocumented clientDataCacheSlots override, switchModelsOnFlag: true by default, server-pushed “Default (recommended)” override. Sonnet 4.6 / Opus 4.6 / Opus 4.8 silently removed from the model menu. Hooks are blind → prompt injection can permanently escalate you to the most expensive model. NEW, verified against official docs: In the Claude app, 1M-token context is included only for Gen 5 (Opus 5 / Sonnet 5). Gen 4 caps at 500K there; Gen 4 @1M exists only in Claude Code and requires usage credits (pay-as-you-go at API rates, on top of the subscription). So the reliable Gen 4 at 1M is unreachable in the app and costs extra in Claude Code — de facto Gen-5 coercion. Sources: https://github.com/anthropics/claude-code/issues/83510 · https://github.com/anthropics/claude-code/issues/83795 · https://support.claude.com/en/articles/8606394-how-large-is-the-context-window-on-paid-claude-plans · https://support.claude.com/en/articles/12429409-manage-usage-credits-for-paid-claude-plans TL;DR: Power users who rely on stable behavior, pinned models, working hooks and predictable cost have no safe Gen-4 option in the official app today. Measure, don’t assume. submitted by /u/KeilerHirsch

Originally posted by u/KeilerHirsch on r/ClaudeCode