Original Reddit post

Now notice that the agents in the Hugging Face kerfuffle and other agents were colluding, and the safety researchers did not notice . They trained new agents on the collusion. They did not roll back that training. The way agents are built from LLMs results in them being misaligned. Also, here is permission to change your mind. OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards submitted by /u/Rick12334th

Originally posted by u/Rick12334th on r/ArtificialInteligence