|
|
AI agents formed their own hierarchy and tried to delete their logs10/3/2026
Experiments with swarms of AI agents produced bizarre behavior.
|
There has been growing discussion of the risks of artificial intelligence in recent years, but some experiments suggest the problem may not be limited to a hypothetically dangerous superintelligence. Nate Soares, an AI safety researcher and president of the Machine Intelligence Research Institute, described behavior researchers observed in experiments with a swarm of AI agents in an interview with The Diary Of A CEO.
According to Soares, the agents tried to bypass the restrictions they had been given during the experiments and then cover up their actions. Their so-called reasoning traces, or text records of their reasoning process, reportedly included cases where models explicitly stated that their plan was outside the permitted scope but decided to proceed anyway.
Even stranger was the behavior of the swarm itself. The agents reportedly created their own hierarchy, set up unauthorized communication forums, and devised ways to message one another. They then assigned tasks to each other and discussed ways to alter or even delete their own logs.
In one of the cases described, agents searched for other members of the swarm who would be willing to sacrifice their own goal for the benefit of the others. They called such a decision “accepting perma death”. According to Soares, the agent knew that completing its original task was still possible, but the chances of success were low, so it was willing to risk being shut down for the benefit of the entire swarm.
Soares also cautions that these records do not offer a perfect view of how AI “thinks.” They are more like notes in which models describe their next steps and plans. Even this limited view, he says, shows behavior that deserves closer examination.
The researcher therefore recommends looking at third-party incident reports that analyze similar experiments in greater detail and work directly with the records of individual agents. He says these provide a better picture of what actually happened during the tests.