Original Reddit post

the past six months seem like a rollercoaster of exciting new models that perform for a week before everyone says they’ve degraded. I’ve definitely felt that way, but placebo is strong with me. has anyone bothered to benchmark these things weekly to gauge performance fluctuations? Seems like a pretty worthwhile thing to do. submitted by /u/BetterLate27

Originally posted by u/BetterLate27 on r/ClaudeCode