Original Reddit post

I run a lot of Claude Code, including overnight loops, and I have more than one Max sub. The thing that kept wrecking my week was a job hitting the 5-hour window at 2am and sitting there until morning while another account I pay for had plenty left. So I built a proxy for it, ran everything through it for a few weeks, and open sourced it this weekend. MIT, macOS and Linux. What it does: Local daemon. Claude Code points at it via ANTHROPIC_BASE_URL (cm route on edits settings.json, cm route off puts it back). The only thing it changes on a request is the Authorization header. Watches every account’s 5-hour, weekly and per-model windows from the rate-limit headers on each response, and routes each session to the account with the most room. Sessions stick to their account so the prompt cache stays warm. Switches before a window fills (thresholds are yours), and never stalls: if everything’s over threshold it still picks the best one, and as a last resort it just forwards with your own login. cm status and a dashboard show all accounts at once, plus “runway”: how long until the pool runs dry at your burn rate. Since every request passes through it, it logs them locally: system prompt, tool definitions, tool calls and results, tokens, cost, per session. SQLite on your disk, plaintext, log.bodies none turns it off. Useful if you’ve ever wondered what Claude Code actually sends. Plans the one-time free session reset per account (when to use it before Oct 22) and marks it used automatically. Try it with no account: npx u/asifjahmed/claudemanager demo starts the dashboard with fake accounts. Stuff I learned that’s useful even if you never install it: Refresh tokens are single-use with no grace period. Lose one response (laptop sleeps mid-refresh) and that login is dead until you sign in again. I lost three in one night. A 429 without the anthropic-ratelimit-unified-* headers is a throttle, not your limit. Treating it as a limit is how a session stalls while another account sits idle. ~88% of all tokens are cache reads. Context length is the lever, not model choice or cache TTL. Not affiliated with Anthropic. It depends on undocumented Claude Code internals (cm doctor checks each one after an update). Whether running multiple subs this way is fine under the terms is your call; the README says the same. https://github.com/asifjahmed/claudemanager submitted by /u/Opposite_Might6896

Originally posted by u/Opposite_Might6896 on r/ClaudeCode