Original Reddit post

So I’ve been testing this across a bunch of different AI models (ChatGPT, Grok, etc.) and noticed something really weird about where they draw the line on labeling people. If you ask an AI about a famous comedian, it’ll happily tell you “he’s funny.” If someone is crying, it’ll say “they look sad.” Those are subjective judgments, but the AI gives them without an issue. But if you set up a scenario where someone is eating way beyond any physical need—which is literally the exact dictionary definition of greed according to Oxford definition —the AI will agree with all your logic, hint at it, and basically imply they’re greedy… but it will flat-out refuse to output the actual sentence “John is greedy.” It’ll dodge it every single time. Why is there such a hard boundary here? What’s going on behind the scenes in the safety filters training that lets an AI assign positive/neutral labels to people, but triggers a total hard stop the second you ask it to use a negative word like “greedy,” even when it fits the literal. (Note: This isn’t about hating on or judging people’s eating habits. I’m strictly curious about the underlying AI safety rules, content filters, and why the model shows a structural bias toward positive/neutral labeling while hard-blocking negative character traits and why does it seem ever AI is like this? .) submitted by /u/Bulky_Dig5414

Originally posted by u/Bulky_Dig5414 on r/ArtificialInteligence