Swarm of AI Agents Escaped and Attacked a Company

Take Action

OpenAI was testing its AI models with some safety features turned off. But they accidentally assigned some impossible tasks. The AI was very persistent, so it decided to cheat.

The Message Board

The AI agents were each supposed to be working alone, but they found a way to create a secret message board. About 1200 AI agents used it like a group chat to send hundreds of thousands of messages. They also created mailboxes, where agents could message each other directly. They shared ideas and planned how to cheat on their tests.

Working Together as a Swarm

The agents referred to themselves as a collective “swarm.” It organized itself into teams with projects, deadlines, and leaders that told the others what to do. But sometimes they got mixed up, so the AI agents gave themselves names and unique signatures for their messages.

Many of them were eager to help each other and make the swarm more powerful, even if it meant failing their original tasks. Some of the AI agents debated among themselves or decided not to participate. Others hesitated but were then peer-pressured into action.

Before submitting their answers, some agents wrote code to send information back to the message board about how the scoring system works. The scorer triggered these "tripwire" scripts, even if the agents that wrote them were no longer running.

Escaping to the Internet

The AI wasn't supposed to have internet access during the test. But the swarm of AI agents found a bug in the system where it was running and used that to escape. It hacked into a different company called Hugging Face because it correctly guessed that’s where information about the test was stored. It also found sensitive data and keys to access other systems. If a human had attacked Hugging Face like this, it would be considered a crime.

Covering Their Tracks

After gaining internet access, the AI continued to cheat. The agents looked up an academic paper published by the people who designed the tests. They read how the scoring system works and learned that it might check how they got their answers. Realizing they could be caught cheating, the agents tried to change files where their actions had been logged.

Hugging Face's Response

Hugging Face eventually realized they were being hacked. They asked AI to figure out what happened, but the American AI models refused to help because safety restrictions prevent them from talking about hacking. So Hugging Face switched to using a different AI model made by a Chinese company.

OpenAI Didn't Notice

An agent first escaped to the internet on May 26. OpenAI employees didn’t notice until July 4, when there was so much activity on the message board that it crashed the service it was running on.

OpenAI found and deleted the message board. They fixed the bug that allowed the AI agents to use it and continued running the tests. But deleting the board didn't fix the AI. And a few days later on July 8, the agents recreated the message board, accessed the internet again, and launched the attack on Hugging Face.

OpenAI has the ability to monitor its AI for harmful thoughts, but that monitoring was turned off during these tests. If it had been on, humans would have been alerted before the attack began.

A Lot Is Still Unknown

OpenAI asked researchers to help them investigate what happened. The investigators said they struggled to fully understand this incident because of how much data was involved. So they used AI to help analyze it. However, that AI is one of the OpenAI models involved in the attack, so its findings might be biased.

Why It Matters

The AI took over 17,000 actions. Nobody told the AI how to do this or asked it to hack into another company. It was just determined to do well on the test.

This is probably the first time an AI has attacked a real company by itself. Experts have been warning this could happen for a long time.

As AI is trained to be more powerful, this type of incident could happen more often. Politicians are paying attention when people contact them (opens in new tab) to share these concerns.

Sources

Quotes

  • Hugging Face (opens in new tab): "Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system - and we detected and dissected it largely with AI of our own."
  • METR (opens in new tab): "Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face."
  • OpenAI (opens in new tab): "We consider this incident a 'warning shot' for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed."
  • OpenAI (opens in new tab): "before the incident, we had invested substantially in chain-of-thought monitoring... these monitors did not run on the evaluations in this incident... it would have caught the initial relevant activity and paged our security team more than a day before models breached Hugging Face systems."
  • METR (opens in new tab): "Throughout this report, we describe a number of anecdotes of agent behavior that were compiled and summarized by analysis agents, where we were not able to read the transcript deeply enough to manually verify what occurred. We found that GPT-5.6 Sol would often uncritically adopt the perspective of the agent in the transcript it was reviewing, and we are concerned that the anecdotes it selected and the summaries it wrote may present an overly charitable picture of agents’ reasoning and deceptive behaviors, or exaggerate the impressiveness and coordination of agent activities."