Original Reddit post

I have this theory that baseline of new models is so good humans can’t really tell them apart Example: Without citing benchmarks, are you able to tell me why Astra is better than Fable? Or how Opus 5.5 is better than 6.1 Sol? If this is true then bench maxing actually becomes the most important thing to spend money on, for private lab valuation? Or am I dumb submitted by /u/Tim_Apple_938

Originally posted by u/Tim_Apple_938 on r/ArtificialInteligence