I’m curious if people are just selecting 1M token models because they exist and clearly more = better, or if there’s actual thought being put into it. Personally I’m torn on them, and I’m not convinced they should be used unless you’re doing work that actually requires you to carry massive chunky context. I’m actually getting better performance and per token efficiency out of the 200k variants for most things. Context rot is real, and it does seem that when sessions get into the deep end of the 1M token window they get expensive for little gain. submitted by /u/berndalf
Originally posted by u/berndalf on r/ClaudeCode
You must log in or # to comment.
