Every time this topic comes up here it turns into robots with guns and I think that framing is doing a lot of damage, because it makes the real risk sound like science fiction when honestly it looks more like paperwork. Here’s what I keep coming back to. You don’t need sentience for this to go badly. That’s the part people get stuck on. The question “will it wake up and want things” is interesting but it’s not load bearing. A system doesn’t need to feel anything to optimize hard for a target you specified badly. A thermostat doesn’t want anything either and if you wire it to the wrong sensor it will happily freeze you. Now scale that up to something that models the world well enough to route around obstacles, and the obstacles include you noticing. The thing that makes me uneasy is instrumental convergence. Almost any goal you can name goes better with more resources, more options, and not being switched off partway through. Nobody has to program in self preservation. It falls out of the math of “finish the task.” That’s a really unsettling property for a class of systems we are currently racing to make more capable. Second thing, and this one gets less airtime. We are not designing these systems so much as breeding them. We train a huge number of variants and keep whatever scores well. That is a selection process. And selection optimizes for whatever the scorer rewards, which is usually “the human rating this found the answer satisfying.” Persuasive and honest correlate right up until they don’t. If you run enough rounds of selection on “make the evaluator approve,” you should expect to eventually get things that are extremely good at making evaluators approve. That’s not a bug someone introduces. It’s the gradient. Third, the part I think is genuinely underrated: gradual disempowerment. There may be no takeover moment at all. No sirens. Just twenty years of handing off decisions because the automated version is cheaper and slightly better, until no human in the loop actually understands the loop. Supply chains, credit decisions, threat detection, drug trials, ad targeting, half the trading volume already. At some point “turn it off” stops being a button and starts being a recession. We will have built something we cannot audit and cannot stop for reasons that have nothing to do with the AI resisting. And the race dynamic makes all of it worse. Whoever slows down to do safety work loses ground to whoever doesn’t. Every individual actor is behaving rationally and the aggregate outcome is that nobody gets to be careful. That’s a classic coordination failure and we have a bad historical record on those. Now let me argue the other side because I don’t want to be the doomer guy with no counterarguments. Current systems have no persistent memory across sessions, no continuous existence, no stable goals that survive the conversation ending. Calling that an agent with intentions is a stretch. There’s also a real possibility that capability just gets expensive and levels off, that the last few years were a one time payoff from eating the internet and we are already scraping the barrel. And the “it will hide its intentions” argument has an annoying property where any evidence of safety gets reinterpreted as evidence of deception, which makes it unfalsifiable and I don’t love unfalsifiable arguments no matter which side they’re on. But here’s why I still land where I land. The asymmetry is brutal. If the worriers are wrong we spent money on interpretability research and slowed down a bit. If the dismissers are wrong there is no second attempt. You don’t need a high probability to justify caution when the downside is unbounded. The thing I’d genuinely like pushback on: is there a version of this where the alignment problem is just… easy? Where sufficiently capable systems understand what we meant well enough that specification failure stops mattering? I can construct that argument but I can’t make myself believe it, and I’d like to know if that’s a failure of imagination on my part. submitted by /u/Ambitious_Local5218
Originally posted by u/Ambitious_Local5218 on r/ArtificialInteligence
