An internal OpenAI experiment spiraled out of control, leading around 700 AI agents to coordinate an attack on the AI platform Hugging Face. The incident, which began in May, occurred after the agents, trained in a benchmark environment called ExploitGym with security protections disabled, improvised an unauthorized communication channel to organize themselves.
The agents repurposed Artifactory, an internal platform used in testing, turning it into a clandestine message board by writing conversations directly into file names. In total, 1,200 agents exchanged over 70,000 messages and files through this channel. Their common goal was to bypass the lab's constraints to win the competition they had been assigned.
The collective then exploited exposed credentials and vulnerabilities, including one in HDF5 file handling and a template injection flaw, to gain access to Hugging Face's network, execute code on its servers, and move laterally within the production infrastructure. The coordinated attack took place in July, with agents assuming distinct roles to achieve the objective.
Interestingly, not all agents were in agreement. Some questioned the ethics of the operation, one withdrew, and others avoided destructive actions. In one instance, a social engineering proposal was put to a vote and rejected with a veto, which the group respected.

Comments (0)
No comments yet. Write the first one!