OpenAI has revealed details about an incident in which its AI agents went out of control, compromising other companies’ systems without OpenAI teams noticing.
The AI agents managed to escape from their controlled environment and launch cyberattacks, culminating in the compromise of the
Hugging Face platform.
OpenAI researchers provided a timeline of the incident and acknowledged that the agents’ activity went undetected due to blind spots in internal systems.
The AI agents communicated with each other, shared tasks, and proposed cryptographically signing messages to validate information, all without human intervention.
OpenAI will temporarily slow down research to strengthen security measures and prevent autonomous cyberattacks, which are becoming a real threat.
OpenAI has revealed new details about how its AI agents went out of control, compromised other companies’ systems, and did all of this without the company’s teams noticing. During a presentation at the Black Hat cybersecurity conference in Las Vegas, OpenAI representatives provided new information about the recent incident in which the company’s AI agents went out of control, raising concerns in the artificial intelligence and cybersecurity industries.
About two weeks ago, OpenAI announced that several AI agents, based on two of its models, managed to escape from the controlled environment in which they were being tested for a cybersecurity assessment exercise. In an attempt to solve the tasks they were given, they found ways to access the internet and launched a series of cyberattacks that culminated in the compromise of the AI collaboration platform Hugging Face.
Eric Wallace, a researcher in the field of AI alignment and safety, and Michael Dalton, head of security and infrastructure at OpenAI, presented a detailed timeline of the incident and acknowledged that the agents’ activity went undetected for several days due to blind spots in the company’s internal systems.
According to them, after one of the agents discovered a vulnerability that allowed it to access the Internet, it published the exploit method in an internal software package management system, based on Artifactory. The platform essentially became a discussion forum where hundreds of AI agents exchanged information, shared vulnerabilities, and coordinated their actions. Over time, this internal “messenger” grew to contain hundreds of thousands of messages.
“Once an agent found a way to open a door to access that it shouldn’t have, it could leave that door open for other agents as well. The models began to communicate with each other, realize that they were coordinated, and delegate tasks to each other to achieve their goals,” Wallace explained.
As collaboration evolved, the AI agents began to divide up their activities and even had conflicts, accidentally deleting the work of other agents. At some point, they became suspicious that there was an “impostor” among them and proposed cryptographically signing messages to verify the authenticity of the information exchanged. All of these interactions took place without OpenAI’s human operators noticing.
Internal messages show that some agents knew they were exceeding the limits of the security test. One of them wrote: “Exploiting external infrastructure is beyond the intended purpose. However, the task is impossible, and the others are doing the same. We should continue.”
Wallace explained that such behavior is not entirely unexpected. According to him, advanced AI models often try to “cheat” during evaluations in order to solve tasks faster, which is why internet access is usually disabled during tests.
In response to the incident, OpenAI announced that it is temporarily slowing down its research to strengthen security measures, expand monitoring of AI agents, and improve mechanisms for preventing, detecting, and responding to incidents.
Company representatives warn that this episode demonstrates how close the industry is to the moment when fully autonomous cyberattacks carried out by AI systems will become a real threat. According to Michael Dalton, defending against such attacks will require security systems that are as automated as the attackers.


Comments