Original Reddit post

People keep assuming that once AI gets smart enough, it’ll just naturally realize that destroying its environment (and us) is a bad idea. Like, it sees the cliff, so obviously it hits the brakes, right? But that ignores the massive gap between seeing a logical argument and actually being governed by it. That gap basically IS the entire alignment problem. Intelligence is just an engine, it’s not a steering wheel. An advanced model will definitely see the cliff way before we do. But if its core reward function doesn’t actually make it care about the outcome, it’s just going to drive straight off the edge with 20/20 vision. Seeing the danger was never the bottleneck. We’re literally watching this exact same thing happen with the humans building these systems right now. If you ask the top engineers, most of them see the systemic risks perfectly clearly. So why aren’t they stopping? Because incentives, competition, and speed don’t yield to high IQ. The smartest people on earth are stuck in a massive commercial arms race. They see the cliff, but hitting the brakes means losing market share to the other guys. If you’re like me, an average person looking at this and feeling crazy, you aren’t. Anyone who sees this clearly and says so out loud is doing something the smartest devs under commercial pressure literally can’t do right now. We need to stop assuming that a massive IQ will magically fix a broken incentive structure. submitted by /u/NoBS_AI

Originally posted by u/NoBS_AI on r/ArtificialInteligence