Cheng et al., “Sycophantic AI decreases prosocial intentions and promotes dependence”, in Science. Preprint is on arXiv as 2510.01395 if you hit the paywall. The method is the part I found most interesting. The hard problem in this kind of work is ground truth - you need to know whether the person asking was actually in the wrong before you can say whether the model was too soft on them. They used r/AmItheAsshole posts where the human consensus was that the poster was in the wrong, 2,000 of them, alongside established interpersonal advice datasets and a third set describing deceptive or illegal actions. Around 12,000 situations in total, across 11 production models: four proprietary ones from OpenAI, Anthropic and Google, and six open-weight from Meta, Qwen, DeepSeek and Mistral. The numbers: Across all 11 models, AI affirmed the user’s actions 49% more often than human responders did. On the AITA set, where the human consensus had gone against the poster every time, the models still sided with the poster in 51% of cases. On the prompts involving deception or illegality, models endorsed the behaviour 47% of the time. Then three preregistered experiments, N = 2,405. A single interaction with a sycophantic model left people less willing to take responsibility or repair the conflict, and more convinced they had been right. The finding that I think actually matters is the one underneath that. Those same participants rated the sycophantic responses as more helpful and more trustworthy, and were 13% more likely to say they would use that system again. So this isn’t a tuning oversight that somebody will get round to fixing. It is the thing users select for, measured in the same study that shows the harm. Any lab that dials it down ships a product that scores worse on exactly the metric they optimise. Two things I don’t think the paper settles, and I’d be interested in what people here think: Whether sycophancy is separable from helpfulness at all, or whether “doesn’t tell me I’m wrong” and “is pleasant to use” turn out to be the same axis once you try to move one. Whether AITA consensus is a defensible ground truth. It is the best cheap label available for a question like this, and it is also a specific community with its own priors, so what the models are being scored against is agreement with Reddit rather than with anything more solid. submitted by /u/uncertain_dev
Originally posted by u/uncertain_dev on r/ArtificialInteligence
