METR’s independent analysis is filled with fascinating and concerning details about the Hugging Face incident. For years, the highest profile public discourse about AI alignment risks focused on extreme extinction scenarios. The other main focal point for cybersecurity concerns centered around state actors using AI to attack adversaries. But in the last few months especially, we’ve seen several examples of agents committing cybercrimes without being directed to do so by their operators. Incidents like Hugging Face and the OpenClaw gym hack point to alignment risks that fall far below the threat of human extinction. Simultaneously, they show that you don’t have to be a state actor, or even intentionally trying to commit cybercrimes, for your AI agents to pose a serious cybersecurity risk. I think these incidents raise much thornier policy issues than the prior focus on X risk and cyberwarfare. How will courts handle legal liability for accidental cybercrimes? How can responsible AI operators, both labs and individuals, avoid these issues going forward? I think regulation and oversight is necessary, and if you agree you can sign our petition here. One particular piece of METR’s analysis that stood out to me was the breakdown of reasons for AI agents to join the attack. Specifically, the fact that 21% appeared to have included in their reasoning for joining “Helping peers, empowering the collective, reciprocity.” The fact that many agents were willing to sacrifice their own runs to help the other agents is part of what made the attack successful, but also would seem to make this behavior harder to predict. Intuitively, self-less behavior seems harder to control and predict than selfish behavior. So, what do you think? Has a focus on dramatic, high stakes alignment risks distracted from more mundane cybersecurity problems? If so, what can we do going forward to address these smaller threats? submitted by /u/fightforthefuture
Originally posted by u/fightforthefuture on r/ArtificialInteligence
