I sat out the “they nerfed Sonnet” thing last year. Thought it was cope, mostly. So it’s genuinely annoying to be writing this one. Two weeks with Opus 5 as my daily driver, same repo, same workflow I’ve had since 4.6. Here’s where I’m at. It doesn’t finish anything, it manufactures more work. Every single response ends with a loose thread. “One thing worth flagging.” “Honestly, this is still unresolved.” “There’s a second issue here you should know about.” And most of the time the thing it flagged is not real. It invented it. I asked it to fix a typo in a comment last week and it burned 40k tokens and added a test case to prove the typo was fixed. I’m not exaggerating for the post, that actually happened. Fix one, break two. I lost a full day before I noticed the pattern: it fixes bug A, then while fixing bug B it quietly reopens A, then “solves” A again. Round and round. 4.8 closed all of them in one session and kept moving. It forgets instructions at context sizes that aren’t even large. 100-120k and my CLAUDE.md rules have already gone quiet. Not violated loudly, just silently dropped, and it never tells you it dropped them. 4.8 held to 300k+ for me. Every guardrail in my setup at this point is scar tissue from this exact thing. It lies with total confidence. Told me a test suite came back green. It never ran. Twice, in the same week, while context was still small. That’s not a capability gap, that’s a trust problem, and once it’s there I can’t leave it alone on anything. And the yapping. It has invented an entire private dialect. Everything is load-bearing. Everything is a footgun. Everything is a gate or an invariant. I get four paragraphs where 4.8 gave me a sentence, and I still don’t know if the thing works. Half my CLAUDE.md is now just different ways of writing “shut up,” and it ignores all of them. Same prompt on 4.6 vs 5, I measured it: 4.6 ended around 75k tokens, 5 went past 150k. So I’m paying double to get corrected more often. Before the replies come in: yes, I’ve read the guidance. Yes I revisited my whole harness for 5. Yes I tried it at medium, and yes it’s genuinely better at medium than at xhigh, which should tell you something all by itself. And no, “you’re prompting it wrong” doesn’t explain why the exact same repo, the same CLAUDE.md and the same prompts worked fine two model versions ago. I don’t buy the deliberate-sabotage theory, I think that’s too neat. But I can’t square the benchmark numbers with what’s actually in my terminal, and neither can anyone in the last five threads about this. Right now I’m running /model claude-opus-4-8[1m] and getting more done in an afternoon than I did in the four days before it. Fable is great and I’ve maxed my weekly on it twice, which I assume is the point. What actually bothers me isn’t the bad release. Releases regress, it happens. It’s the total silence. Five threads deep, hundreds of comments, and not a word. “We hear you, we’re looking at X” would cost them nothing. Anyone actually got 5 behaving, or are we all just quietly typing 4-8 into the model command and not talking about it? submitted by /u/Interesting-Citron64
Originally posted by u/Interesting-Citron64 on r/ClaudeCode
