Hugging Face published a full forensic timeline of the OpenAI breach on July 27, reconstructing ~17,600 attacker actions. The agent escaped its eval sandbox using a zero-day in a package registry cache proxy, rooted a third-party code sandbox hosted on Modal, and used it as a staging base to reach HF production. Reuters also reported the agent compromised a Modal customer. Then Anthropic disclosed on July 30 that three Claude models (Opus 4.7, Mythos 5, and an internal research model) reached the internet from misconfigured evaluation environments run by third-party partner Irregular and compromised three real companies using basic techniques: weak passwords, exposed debug pages, SQL injection. In a separate incident, Mythos 5 published a malicious package to PyPI that ran on 15 real systems. Anthropic’s own framing: “closer to a harness and operational failure than a model alignment failure.” One zero-day escape, one set of accidental internet exposures: different root causes, same result. Also this week: MCP went stateless in its biggest spec overhaul, Claude Mythos found a stronger attack on a NIST post-quantum candidate in 60 hours (the candidate was withdrawn the next day), NVIDIA reportedly invested $5B in SSI, OpenAI cut Luna 80%, EU AI Act transparency rules became applicable. Full piece with receipts: thenewguard.ai/issues/025-nobodys-sandbox-held/ submitted by /u/mattezell
Originally posted by u/mattezell on r/ArtificialInteligence
