“Arena ai”/LMArena were probably the first to transform human preference data into a well known and highly important metric for LLM providers. However, there were voices and blogposts e.g. from the people at surge ai that say that arena is " cancer on ai ". The arguments are quite convincing, and i suppose most model providers use their rank on Arena only as marketing tool anyways - to not reproduce the sycophancy crisis. The people at Max Planck Institute for Intelligent systems, at the social foundations of computation department do a lot of benchmarking research and recently published " comparity.ai ", which in spirit is similar to arena, but has some different quirks. Whats i really like is the idea behind the personal leaderboard, which updates with your votes, so after playing around there a little bit, you see what models work best for you (so after that you can actually say that e.g. ChatGPT is better/worse for you than Claude) submitted by /u/adam_alpha_finetuner
Originally posted by u/adam_alpha_finetuner on r/ArtificialInteligence
