Original Reddit post

I have a fairly complex platform, main infrastructure on one GitHub private repo, second repo has all the accumulated data, my data moat, so to speak. I am using a very simple setup: I discuss optimal, efficient architecting with my other AI tools, and I send a comprehensive relay with scope/brief for execution to Claude Code CLI, burning API tokens, on my Mac Terminal . I am using an Anthropic API key, not Max, so that I dont run into any weekly limits with a Max plan. I had previously been using Opus 4.8 on my CLI, but I downgraded to Sonnet 5 which is cheaper in order to save on API costs. I had no issues with good execution at Sonnet 5 level. But lately, I have had huge token spend which I can track on my Claude Cost Console. I had Claude Code do an anlysis of one sprint, LOC changes, and actual actions. The total spend on the sprint was 132K tokens. Start to Finish including merge to Main was only 26 minutes, unbelievable, with zero errors. No way one can minimize that aspect. Real numbers, from git diff across the whole sprint: Platform repo: 43 files changed, 419 insertions / 234 deletions (~650 lines touched) Data repo: 8 files changed, 690 insertions / 2 deletions (~690 lines touched) Total: ~51 files, ~1,340 lines touched So the actual code volume was modest — this wasn’t a huge refactor by line count. The token spend came from the shape of the work, not the size of it. A few concrete drivers: Fan-out across many small touch points, not one big change. Adding “X” meant editing the same handful of facts in 3 components × 4 languages, plus the 8→9 sweep touched ~40 unrelated files each needing a 1-2 line fix. That’s dozens of separate Edit calls for a few hundred net lines — expensive per line changed compared to one contained refactor. Repeated npm run build runs. Each one prints the full Next.js route list (100+ lines), and I ran it manually several times per round plus Husky’s pre-commit hook re-runs it automatically on every git commit. That’s real, avoidable duplication — I was effectively paying for the same build output twice per round. Verification research. I didn’t just take Cowork’s word for facts — I independently searched the web for “Y”, etc. Each WebSearch call returns fairly verbose snippets, and I ran several per round across three rounds. Three separate rounds, not one. The initial sweep, the “Z1” round, and the “Z2” round each repeated the same machinery — build, commit, push, verify live, write a relay — rather than being one pass. Those of you who are experienced CLI users, what can I do to reduce spend here? Is repeated npm run build an issue, can I avoid it without getting into trouble? Is starting a new session on the Terminal important? I don’t think I do it enough. What can I do to reduce simple web surfing search costs? I would deeply appreciate any help. submitted by /u/OneHuman_aiprotect

Originally posted by u/OneHuman_aiprotect on r/ClaudeCode