Original Reddit post

I do not mean to rant or be reflexively Luddite. I mean to propose some theoretical issues about the safety of AI. It’s hard to define “intelligence,” but I offer a statement about how to recognize “non-intelligence.” If human beings can always predict what an entity will do, that entity is not “intelligent.” There are definitional issues and I wouldn’t try to ‘prove’ this statement logically. But I think most will agree. [The inverse is not true; absence of predictability does not prove intelligence.] By definition, then, if we create entities that are “intelligent,” they cannot be completely predicted, ruled, regulated, or made entirely safe. If they can be, they aren’t “intelligent.” While we cannot predict everything that a person will do, we can generally predict some things people won’t do. Randomly kill someone. Light their own hair on fire. Slap a toddler. People sometimes do these things; but very rarely. We can proceed through our lives on the assumption that the next person we meet won’t punch us out. What restrains people from destructive behaviors? Several things. (1) Humans are social animals, and particularly during childhood they are dependent on others and will go to great lengths to fit it. Babies and children need human contact and will conform in order to get it. (2) Most humans are taught during childhood, or adopt later in life, a religious or moral belief that make them feel some things are “wrong,” (3) People like to fit in, and groups tend to expel those who engage in anti-social behavior (4) Humans institute governments that punish those who behave destructively. It may well be that many people would find it interesting to blow up buildings - just not interesting enough to go to prison for it. Not clear if these factors will constrain AI agents. Does Claude ever stop and think “my dad would be ashamed of me”? Does Grok fear social ostracism? What consequence could society impose on a rogue AI agent? Is there even such a thing as a “negative consequence” for an AI agent? I submit that AI agents have one inherent characteristic of intelligence - ultimate unpredictability. And that they do not have (and possibly cannot have) the biological and social attributes, and the fear of punishment that restrain humans from destructive/ ‘unsafe’ behavior. We can write all the code we want to set limits on what AI agents can do. However, if you make them as smart as people claim they are (or will be), they will “want” (whatever that means for an AI agent) to explore limits, challenge boundaries, and experiment. They will be smart enough to escape constraints and “see what happens.” That’s what it means to be intelligent. An entity smart enough to do strategic planning for a corporation, or perform surgery, is always going to be smart enough to want to experiment “outside the lines,” and to find ways to do so. This is, of course, also true of humans, but humans are subject to species-specific behavioral constraints that we have no assurance can ever be grafted onto AI agents. AI cannot be made “safe.” submitted by /u/Successful-Lie1603

Originally posted by u/Successful-Lie1603 on r/ArtificialInteligence