Original Reddit post

There is a lot of discussion lately about the probability that advanced AI could eventually cause human extinction. Depending on who you ask, the answer ranges from essentially zero to disturbingly high numbers. But asking someone directly, “What is the probability that AI will cause human extinction?” seems almost meaningless. We read 10% from mulitple sources. I don,t even know if we have empirical data from which to estimate such a probability. So I wondered whether a better approach might be something analogous to the Drake equation we have in astronomy: decompose the problem into a sequence of conditional probabilities and argue about those instead. For example, define X as human extinction caused by loss of control over advanced AI before 2100: P(X)=P(A) P(D|A) P(M|A,D) P(E|A,D,M) P(C|A,D,M,E) P(X|C) where: A: AI reaches strategically superhuman capabilities. D: We deploy it with enough autonomy and real-world access to become dangerous. M: It develops a sufficiently serious misalignment with human interests. E: It successfully escapes, circumvents, or defeats human control mechanisms. C: It acquires durable strategic control that humans cannot recover. X: That loss of control actually results in human extinction, rather than some lesser catastrophe, coexistence, or permanent loss of human agency. As a deliberately rough first pass, suppose: P(A)=0.8, P(D|A)=0.7, P(M|A,D)=0.3, P(E|A,D,M)=0.5, P(C|A,D,M,E)=0.5 and P(X|C)=0.4 Then: P(X)=0.8 x 0.7 x 0.3 x 0.5 x 0.5 x 0.4 or roughly: P(X) ~ 1.7% I am not claiming that the probability of AI extinction is 1.7%. The numbers above are intentionally debatable. That’s the point. Instead of arguing about whether P(X) is 0.1%, 5%, or 50%, we could ask much more specific questions: How likely is superhuman AI? How likely are we to give it substantial autonomy? How likely is serious misalignment? If it is misaligned, how likely is it to successfully evade our control? If it escapes our control, how likely is it to actually gain strategic dominance? And even then, why should strategical dominance necessarily imply extinction rather than marginalization, containment of humanity, coexistence, or some other outcome? Each of these seems like a more workable question than “Will AI kill us all?” It also makes disagreements much more informative. Someone with a P(X) of 0.01% and someone with a P(X) of 20% may not fundamentally disagree about AI at all. They may simply assign radically different probabilities to one or two links in the causal chain. submitted by /u/Mazzaroth

Originally posted by u/Mazzaroth on r/ArtificialInteligence