Original Reddit post

Feels like I was in shock after Opus 5 output and stopped reading docs. there’s still very good ones being released, including this deep dive into Terminal-Bench 3.0. written by Anthropic employee working in the Claude code team. There’s a lot of interesting stuff about the cache, effort and benchmarks but my core takeaway was we still need to plan and spec well. Here’s 9 key points covered in the doc split into 3 sections: Section 1 - the benchmark data - how Anthropic tested these features and what the numbers mean. Numbers are internal. anthropic’s own runs, 5 tries per task, some safety features off. they won’t match public leaderboards The benchmark tasks are extreme. things like building a game console chip or writing a formal math proof. much harder than normal coding… In theory means better at day to day tasks. In theory Each step up scores better and burns more tokens. changing effort in claude code no longer breaks the prompt cache. Biiiig one imo some fields gain far more. security and hardware jump a lot. rule-heavy admin work barely moves. This is just winning imo Section 2- how effort mechanics work - what is actually happening under the hood when you increase the effort level. a detailed spec makes effort matter less. vague prompt: higher effort means more features and more guesses made for you e.g. fitness app: 1.5 min on low, 67mn on max. full spec: all levels land close. Tip: just plan and spec well effort catches missed edge cases, not a wrong approach. if claude misunderstands the task, more effort won’t save it. what effort is: a time budget. higher means claude tests more, checks its own work, and makes more calls itself. Section 3 - practical workflows - how to apply these settings to your daily development routine. rule of thumb: low for brainstorming and quick edits, medium for everyday features, high for bug fixes in existing code, max for hard fully autonomous jobs… the loop to copy: have claude interview you about your spec, build on low, review and iterate on low, then test and verify on high. switch anytime with /effort Notable quote (the first sentence): “One of the best parts of our newest Claude models is how they respond to effort without breaking the prompt cache in Claude Code” submitted by /u/BuffaloConscious7919

Originally posted by u/BuffaloConscious7919 on r/ClaudeCode