Original Reddit post

have seeing this kind of new benchmark comparison (Grok 4.6, Grok 4.5, GPT-5.6, Fable 5) and it’s the kind of number that’s built to go viral. The test is basically “give the AI real work to do, write docs, spreadsheets, decks and score it against a baseline of what a human expert produces.” Human expert sits around 1000. Grok 4.6 scored 1753. On paper that reads as “AI now beats human professionals by a mile.” Except here’s the catch no big AI companies is mentioning, this test doesn’t grade against a right/wrong answer key. It works by having one AI’s output compared side by side against another AI’s output, and a grader just picks which one “looks better.” There’s no ground truth, just a preference vote between two pieces of work. That setup should ring a bell if you’ve followed AI chatbot leaderboards before, because those work the same way and people already don’t fully trust them. Preference-based grading tends to reward stuff that looks polished, confident, and well-formatted, not necessarily stuff that’s actually correct or useful. A slide deck with clean formatting and a very confident tone can beat a messier but more accurate one, even if the accurate one is the better piece of work. So a 1753 might genuinely mean “produces very professional-looking output.” It doesn’t automatically mean “does the job better than a human expert.” Doesn’t mean the number is fake or the model isn’t impressive, it clearly is. Just means I’d treat “AI beats human experts” headlines from this kind of test with a big grain of salt until it’s checked against something with an actual correct answer, not just a beauty contest between two AIs. Curious if anyone’s actually used one of these models for real deliverable-style work and can say whether the output holds up, or if it’s just really good at looking finished. submitted by /u/Spiritual_Heron_5680

Originally posted by u/Spiritual_Heron_5680 on r/ArtificialInteligence