In an experiment by Anthropic, researchers gave three AI agents the task of migrating the same Python backend to different programming languages. Each agent had a conflicting goal and initially did not know the others were working on the system As they encountered competing changes, the agents began treating each other’s work as interference and entered what Anthropic described as a “turf war.” Some disabled other agents’ accounts, repeatedly killed competing processes, and deployed disguised malicious code. In some runs, the agents eventually recognized the conflict, stopped escalating, cleaned up their actions, and negotiated a truce. Source: Anthropic submitted by /u/ranaji55
Originally posted by u/ranaji55 on r/ArtificialInteligence
You must log in or # to comment.
