Benchmarks look good but I’ve tried using it on multiple occasions in real code and it face plants over and over again. It will find some load-bearing comment in the code and then that becomes the basis for everything else going forward. It might work in isolated environments, but so far I haven’t found it very useful. Maybe this is because it’s designed for short tasks without any context. For real usage I’m still finding that deepseek flash 4.1 is simply better in every way for a cheap model to do simpler tasks. Deepseek flash 4.1 is pretty good and very cheap especially on a go plan, haiku seems cheap but it just burns through tokens falling into load-bearing trap after load-bearing trap. submitted by /u/YearLight
Originally posted by u/YearLight on r/ClaudeCode
