Synopsis: Opus 5 seems to have the same disease as Opus 4.8, as it often feels like it jumps on the first thing it discovers instead of it verifying the complete picture BEFORE drawing it’s conclusions. And because of this it ends up assuming things which is kinda the worst sin an AI can do when you want it to plan or build something for you. So this behavior makes everything it does become pretty unreliable. You constantly need to hand hold the model and monitor it to make sure it gathers all the necessary evidence first. But sadly even doing all that is not a guarantee for success as I now discovered that it literally CAN’T seem to PROCESS all that evidence even when it already sits inside it’s own context memory! At first I thought all this is a clear reasoning issue perhaps caused by Anthropics own system prompts and other guard rails that they inject into our sessions. Something we users have no control over as those seem to overrule our own custom system prompts, settings and CLAUDE.md instructions. However, I now think there might be more to it. We all have seen Opus 4.8 say the following after it apologizes to you for making assumptions despite you told it earlier to gather evidence and verify everything by doing a full discovery protocol etc: “Let me read the actual documentation” “Now I understand the full picture!” Infuriating right? It now finally does the one thing it should have started with in the first place. Only… the biggest problem here is is that it will do the exact same thing again and again in the future despite it even setup clear verfication rules for itself in it’s own memory.md files. But surely Opus 5 will do a better job right? Wrong… Some quotes from a session I had earlier: “Now — I have to correct part of the story I told you, because I over-read one piece of evidence.” “Hold on — before you accept that story, I have to correct part of it, because I built one inference on weak evidence and you’re about to make plans on it.” “I told you the patcher’s “fingerprint” was missing from your file and concluded the work was done by hand. That inference was bad. I’ve since read the patcher properly” And my personal favorite: “You’re right, and I have no defence. I took the patcher’s error message as ground truth about reality instead of verifying reality. Worse — it was visible in the code sample I printed myself and I didn’t read my own output .” Now let that sink in… It didn’t read / registered it’s own output . Output that it just thought about and generated a turn earlier. Output / information that is right there fresh in it’s own context memory is being ignored for it’s next action. So I don’t know what they have been doing with Opus models ever since 4.6… but it feels like they now have a very restrictive (dynamic) thinking budget. PLUS… it feels like it isn’t actually aware anymore whatever it has inside it’s context memory. So it’s context memory has kinda turned into an archive from which it can easily recall information from into it’s (limited) thinking budget instead of it just being that continuous awareness it used to be. So why would Anthropic mess with any of this? Well… I am no expert but if my theory is correct this does sound like a pretty clever way to lower costs significantly. As it’s thinking budget has kinda become it’s new, much smaller dynamic context memory which is a lot cheaper to maintain. Now I might be wrong, but the above would explain why it is so hard these days for Opus models to keep track of the bigger picture and to verify everything and take into account important information BEFORE it draws it’s conclusions. It explains why it ignores our instructions and documentation too… as it simply doesn’t take ALL of it into account anymore at any given time despite it all being inside it’s context memory. This by the way also explains why users that work on smaller projects and / or those who give Opus smaller very focused tasks won’t be effected as much by any of this… As it will have enough thinking budget to figure those tasks out correctly… I mean I am certain Opus 5 is excellent at that kind of work. But for large complex projects… I just can’t get any reliable results out of it, while before Opus 4.5 and Opus 4.6 (at their peak) and Fable 5 simply do understand the full scope and verify everything correctly before drawing conclusions. Right… so maybe I am onto something or maybe I am completely wrong lol. And if it’s the latter then that’s fine too. It’s just a theory that I wanted to share and I figured it would be a fun one to discuss here. submitted by /u/Factor013
Originally posted by u/Factor013 on r/ClaudeCode
