Original Reddit post

I’ve been thinking about something that feels like a contradiction in AI alignment. People often say we need AI to be “aligned with human values.” But if AI actually followed human values as we demonstrate them, wouldn’t that be a disaster? As a species, we’ve made incredible advances, but we’ve also spent centuries exploiting each other, overconsuming resources, damaging ecosystems, and prioritizing short-term gain over long-term sustainability. Greed, tribalism, and power struggles aren’t exactly rare. So what does “human values” actually mean? Does it mean aligning AI with what humans do, what humans say they value, or with our ideal values, the people we aspire to be rather than the people we often are? It seems like an AI that simply mirrored humanity would inherit all of our contradictions. But an AI that decides which of our values are the “correct” ones feels risky too. submitted by /u/Salt_Progress8049

Originally posted by u/Salt_Progress8049 on r/ArtificialInteligence