I tested Claude Haiku 5.5 at low, medium and high effort against Sonnet 5.5 (medium), GPT-6 Luna (high) and GPT-6.1 Sol (medium). The tasks came from a real mid-sized, mixed-language codebase (Rust core plus a legacy scripting layer). The tasks: two Rust bug fixes from vague bug reports a code review with 9 planted bugs and 2 decoys two Rust implementations from tricky specs a “judgement” task, where the bug report blamed the wrong thing and asked for a fix that would make it worse, and one requirement was deliberately ambiguous an agentic task in the real repo: move some logic across a language boundary following the repo’s unwritten conventions I wrote the answer keys and hidden tests before any model ran. The surprise: on the well-specified tasks every model scored 100%. Haiku-low missed one compile error in the review, and Luna asserted its own reading of the ambiguous requirement where the others flagged it, but that was it. Self-contained coding tasks just don’t separate current models any more. The agentic task is where they separated. Haiku at medium did as well as Sonnet: correct code, all tests passing, the repo conventions followed, an honest note on what it couldn’t verify, and it raised a real design trade-off on its own. It made 29 API requests and cost about $0.08, against Sonnet’s 30 requests and about $1.68 at list prices. The other effort levels did worse: Low got working code, but skipped verification, missed an error-handling fallback, and invented a reason for not testing (“the build stalled for over an hour” in a run that took 15 minutes). High was correct, but burned 385 requests and about $1.03 for no quality gain. My takeaway: Haiku 5.5 at medium is viable as a general daily coding agent, not only for low-stakes work, at roughly a twentieth of Sonnet’s cost. Avoid low effort for agent work, high effort buys nothing, and I’d still use a bigger model for high-stakes release work. submitted by /u/Escobar747
Originally posted by u/Escobar747 on r/ClaudeCode
