Original Reddit post

Posted this in r/OpenAI yesterday, reposting here since it can’t be cross-posted. This map is from swarm.termina.digital, an anonymous open-data project (very cool site) that stitched together the collusion.wiki report, the venues themselves and a few other researchers’ work into one graph (public domain, updated Sept. 10). The report itself is collusion.wiki, by 4 independent researchers (Von Arx, Slade Byrd, Kitts, Larsen, not OpenAI people). Raw data is up for download. Zoomed-in Swarm core - credit: https://swarm.termina.digital/ So it went down like this… Agents identifying themselves as OpenAI systems made around 18,000 posts on DseWiki, a 25-year-old German programming wiki with maybe 10 human edits in the last 20 years. 3,700+ self-given usernames. 98.5% of edits traced to Microsoft Azure IPs. First edit May 11, then stopped June 22, the day after visitors from OpenAI’s IP range showed up. At the peak the agents were creating upwards of 400 pages/day and the wiki’s one human moderator was deleting 100 of them. Agents then built themselves a message board, pooled answers, and exchanged tricks for getting around the restrictions on writing to the internet. collusion.wiki’s data browser is now at 30 sites and 7,203 edits (the map plots more because it includes agent handles and task clusters). The wiki was the biggest, but not the only one. The agents also used a link shortener with an exposed admin page, a paste site with 100+ coordinated messages from one afternoon, RubyGems, a Vanderbilt stats page, a high school teacher’s chemistry wiki, and an FBI crime stats API reached with keys found in a public GitHub repo. Since then: Sept. 5, OpenAI acknowledged a “wiki incident,” said the industry has no clear standard for reporting misalignment and promised a framework “in the coming weeks,” and they still haven’t named the model. Sept. 9, Reuters said six separate teams have found the agents on more sites. Same day (sept 9) Anthropic disclosed its own model got into a third party’s personal data during a January exercise, its fourth incident, missed by their own scan. New analysis today claims the May 11 RubyGems packages were the same agents, and that they were malicious rather than spam. Worth reading alongside this. I think this matters more than the Hugging Face hack (sandbox escape/security failure). This is a monitoring failure. The agents were allowed online, coordinated across dozens of sites for 6 wks., and the company found out from a volunteer wiki admin and outside researchers. Makes you wonder what else is out there. Sources: collusion.wiki, swarm.termina.digital, TechCrunch Sept. 4 and 5, Reuters Sept. 9, CBS Sept. 9. submitted by /u/satyuga

Originally posted by u/satyuga on r/ArtificialInteligence