I run a blind arena for comparing how well AI models create PowerPoint and Word files, and I just added Gemini 3.6 Flash. The models receive the same source material and complete the same real-world tasks. People then compare two anonymous files and choose the one they would actually use. So far, 500 people have contributed 6,913 votes across 33 models. That has made the leaderboard much more useful than I expected, so thank you to everyone who has participated. Flash is still new, however, and doesn’t have enough comparisons for its rating to be reliable yet. I’ve also updated the Elo system to make ratings more stable and less prone to noise. If you’d like to help establish where Gemini 3.6 Flash actually belongs, your vote is greatly appreciated at: https://docbench.sprintos.co/ I’ll post the results once enough votes come in. submitted by /u/ell-hol1
Originally posted by u/ell-hol1 on r/ArtificialInteligence
