Hallucination was often spoken of as a major, fundamental, intractable problem of LLMs. Then they started measuring it in a public and standard way and after that the problem rapidly got better. The benchmark was published at the end of last year ( https://arxiv.org/abs/2511.13029 ) and almost every model that scores above 0 has been released after that point, and the general trend is that newer models, especially frontier models, are improving on this benchmark. It’s true that correlation doesn’t equal causation but there are some good reasons to believe the two are in fact connected in this case: 1 - If you don’t measure the hallucination rate, only the percentage of correct answers, then models are heavily incentivised to just guess if they don’t have a high-probability answer. They lose nothing if they’re wrong, compared to refusing to answer. But they might get lucky and be right. Hallucinations were created by the “count only correct answers” benchmarks. If you do measure and punish hallucinations, it can be and has been trained down to much lower levels. 2 - Popular benchmarks are heavily trained for. People dismiss benchmaxxing and while it can be a problem in individual cases, models need to be measured somehow. The broader the set of benchmarks, and the more accurately the individual benchmarks measure what they intend to measure, the smaller the gap between “training to the test”, and actually being good at the tasks in question. What does this mean for control and alignment? For users, it means we need good benchmarks specifically for transparency, honesty, and obedience. That is: does the model try to hide what it’s doing; does it actually (in its tool calls, chain of thought, network usage) do what it reported it was doing; and does it do what you told it to, and not do what you told it not to do. For the other side of alignment: conflicting interests and weaponisation, the problem is not a testing problem, it’s a competition, cost, and openness problem. If models are well aligned, broadly speaking, to their actual user’s interests and instructions, then the safest thing is that as many protectors, defenders, and benign users have access as possible, to help find and fix their their own vulnerabilities, and prepare for what malicious users might try. Having a smaller number of providers, or mostly closed models, is the less secure option on that front. submitted by /u/StrategicHarmony
Originally posted by u/StrategicHarmony on r/ArtificialInteligence
