Not a rant — measured and source-verified. Two open issues consolidate the evidence:
Quality regression (Gen 5 / VCST): BullshitBench nonsense detection Gen 5 0.41–0.74 vs Gen 4.6/4.8 0.83–0.95 (95 % Wilson CIs, no overlap), +84–107 % verbosity at identical reasoning effort, silent rerouting (Fable 5 → Opus 4.8).
Model pinning is broken by design: 4 bypass vectors — /model internal write (0 hook events), undocumented clientDataCacheSlots override, switchModelsOnFlag: true by default, server-pushed “Default (recommended)” override. Sonnet 4.6 / Opus 4.6 / Opus 4.8 silently removed from the model menu. Hooks are blind → prompt injection can permanently escalate you to the most expensive model.
NEW, verified against official docs: In the Claude app, 1M-token context is included only for Gen 5 (Opus 5 / Sonnet 5). Gen 4 caps at 500K there; Gen 4 @1M exists only in Claude Code and requires usage credits (pay-as-you-go at API rates, on top of the subscription). So the reliable Gen 4 at 1M is unreachable in the app and costs extra in Claude Code — de facto Gen-5 coercion.
Sources:
https://github.com/anthropics/claude-code/issues/83510
·
https://github.com/anthropics/claude-code/issues/83795
·
https://support.claude.com/en/articles/8606394-how-large-is-the-context-window-on-paid-claude-plans
·
https://support.claude.com/en/articles/12429409-manage-usage-credits-for-paid-claude-plans
TL;DR: Power users who rely on stable behavior, pinned models, working hooks and predictable cost have no safe Gen-4 option in the official app today. Measure, don’t assume.
submitted by
/u/KeilerHirsch
Originally posted by u/KeilerHirsch on r/ClaudeCode
