I work in a variety of projects. Biotech, stock data analysis, straightforward UI+backend type startup, audio signal processing. The latter project has been the toughest thus far to crack as it’s a legitimately hard problem I’m doing in particular. I spent ~2-3 weeks down in the weeds with: GPT 5.6 sol, extra high. Fable 5, extra high, Opus 4.8 extra high and Opus 5 extra high. I did give them detailed directions on my insights and such but we kept making progress here and there, then finding out “oh those gains weren’t real all along it was leaking information across train/test” or something of that flavor, restarting progress, etc. Then I decided to just give Opus 4.6 a try. I’ve always been partial to it because it: actually listens to you and understands your intent. 4.8 would misinterpret things “ah I must remind you that no that’s wrong” then I clarify “oh I see, I still need to tell you”. Fable 5 sort of just runs 15 steps ahead by itself, doesn’t interpret the results back to me, but I do feel the higher level of intelligence as it picks up on things. does not overengineer things. a problem I really noticed was opus 5, fable 5, possibly gpt 5.6 would overengineer the solution past what I asked. Yes, they’re thinking of more things, but at the same time I find out in domains I’m not as privy to that there’s so many extra things they adopted vs trying out the simplest solution that most closely aligns with what I narrate out. GPT 5.6 would go super, super deep investigating just the one thing and while that’s nice I need to be able to course correct sometimes. Opus 4.6 in contrast tries to implement something really close to what you asked and does the simpler solution first. Lo and behold, a blind experiment with opus 4.6 hit / exceeded what those frontier models spent 2-3 weeks working on with multiple restarts, just because it listened to exactly what I stated based on my understanding of the problem, and we’re iterating at just the right pace. That’s also not to say - the amount of times Opus 5 would mess up on simple things in a fresh context with a well prepared architectural setup with clear paths to data. “Stop. The data was interpreted incorrectly. Good thing you asked for another pass, this was wrong” repeated 2-3 times. Hard to trust it… you can actually read its thoughts This is really great so you can instantly correct any possible mistakes or misunderstandings. For example, it trying to search for what you asked, just tell it the paths and clarify instead of wasting all the output tokens. It’s just faster to be able to stop and course correct earlier on. Well as you just read, why am I working better with a model that’s ~5 months outdated, worse on benchmarks, etc.? I can’t solve everything but most of the experience is pleasant and I can trust it more. Sometimes I wonder if Opus 4.6 was lightning in a bottle and we can’t go back to it. It was just the right amount of “smart” to get what you want to do, narrate stuff back to you in a format you actually understood, when I started getting a bit lazy and shorthanded my thought process it understood perfectly what I meant based on our previous chat history and didn’t assume I’d betray the trading setup or that I misunderstood something. The other models do have their advantages, and have helped me accelerate past certain issues by spotting things earlier models didn’t. But man, do I keep coming back to opus 4.6 and actually making progress with it. Maybe call it a skill issue or whatever but I’d like to think I know what I’m doing after spending 5 months exclusively agentic coding, developing routines to make sure my intent is carried out, etc. I really hope they keep this one around for a lot, lot longer. submitted by /u/thrandomaway
Originally posted by u/thrandomaway on r/ClaudeCode
