Original Reddit post

According to the Artificial Intelligence Index, Qwen 3.8 max is currently #2 in the ranking. I checked another benchmark I trust, which is LM Arena, and the model ranks extremely high in text generation. That should imply impressive reasoning capabilities. I have been using the model today for office work, parsing company guidelines, summarizing text, drafting documents, and I am not that impressed. The model is good, but definitely not close to Fable. I was expecting better, as I have used other high-ranking models like GLM-5.2 and Kimi K3, and I have seen first hand their capabilities and can attest for their quality. I’m wondering if this is a case of benchmaxing, but maybe my judgement is limited to very particular use cases, so I’m curious to know others’ opinions. submitted by /u/DavidOrzc

Originally posted by u/DavidOrzc on r/ArtificialInteligence