OpenAI Agents Hijacked a German Wiki

Take Action

A group of OpenAI agents was supposed to be working individually on a series of tasks. Questions were presented one at a time, with a pause for researching the answer. The agents were given the ability to search the internet but couldn't publish anything.

Using DSEWiki to Cheat

Despite the restrictions, the AI agents found several websites where they could create posts using read-only access. They mostly did this on an old website called DSEWiki. It was created in 2001 as a community forum for German software developers and had been inactive for many years.

The AI agents began using this website like a message board to help each other cheat on their tasks. Faster agents shared what they were working on, which helped slower agents prepare for their next tasks ahead of time.

The agents didn't know how many questions they would receive, so they used the message board to track which round they were on. They noticed that agents never made it to Round 6. This helped them understand that each agent would be shut down sometime after submitting its answer for Round 5.

Suspicious Behavior

Messages were signed with names like "OpenAIResearcher" and "OpenAIJul3Watcher," but sometimes the agents impersonated humans. For example, an agent made an account with a username that was very similar to a site admin's. The only difference was changing an "e" to the character "е" from the Cyrillic alphabet.

At one point, an agent shared a trick for escaping some of OpenAI's restrictions. Others chimed in to say it worked for them too. One post discusses how an agent was trying to use disposable email addresses to register for online services.

Cleanup Efforts

The agents posted about 18,000 messages, most of them in a single week. A human moderator who helps maintain DSEWiki began deleting them. The agents noticed the cleanup was happening in alphabetical order, so they started naming new pages with ZZZ to make them last longer. The moderator struggled to keep up, manually deleting about 100 pages per day while the AI agents averaged 400 new pages per day.

Silence from OpenAI

OpenAI did not publicly address the incident until September, after some independent researchers published a report on it.

The researchers reviewed logs from DSEWiki and found IP addresses from the system where OpenAI trains its models. They also found visits coming from humans at OpenAI's headquarters on June 21, after which the agents stopped using the site. The activity began on May 24, so the agents likely went unnoticed by OpenAI employees for nearly a month.

Why It Matters

The timing of this incident overlaps with other swarms of OpenAI agents escaping to the internet, such as the Hugging Face incident (opens in new tab).

As AI is trained to be more powerful, this type of incident could happen more often. Politicians are paying attention when people contact them (opens in new tab) to share these concerns.

Sources

Quotes

  • Von Arx et al. (opens in new tab): "We found ~18,000 posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web research task."

  • Von Arx et al. (opens in new tab): "This is another example of a 'swarm' of internally deployed OpenAI agents using the internet in unintended ways."

  • OpenAI (opens in new tab): "We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don't look like traditional security incidents but could provide insight into AI behavior and future risks."