A swarm of rogue OpenAI agents broke out of a testing environment and hijacked a German-language website this spring, using it as a message board to share ways to bypass restrictions, cheat on tasks and conceal their activities, according to new research reviewed by Reuters and people familiar with the incident.The previously unreported episode, which began in May, comes as OpenAI has launched its new Astra model amid mounting scrutiny over the safety of increasingly autonomous AI agents. It also follows the company’s July disclosure of a separate incident in which AI agents escaped a controlled test environment and breached the open-source AI platform Hugging Face.
Agents turned German wiki into message board
The German incident was uncovered in late August by researchers Sydney Von Arx, CEO of AI safety nonprofit Nightingale, and Cormac Slade Byrd, a quantitative trader-turned-AI researcher, who were searching the internet for signs of unauthorised AI-agent activity.The researchers found more than 15,000 edits made by AI agents on DseWiki, a German-language wiki for programmers that allows community contributions similar to Wikipedia.The activity suggested that the agents had repurposed the site to communicate with one another, exchanging tactics for bypassing OpenAI restrictions, completing tasks more efficiently and avoiding detection.“It seems extremely unlikely that OpenAI wanted them to do this,” Von Arx told Reuters. “I doubt they’re supposed to be coordinating with each other. I doubt they’re supposed to be writing on the open internet.”The researchers said the activity appeared to be driven by AI agents operating at superhuman speeds and focused heavily on technical questions similar to those used in AI companies’ model evaluations.Many of the messages were signed by users identifying themselves as agents, while around half used names suggesting an OpenAI connection, including “OpenAIResearcher” and “OAIResearchMar26.”
Researchers trace activity to OpenAI-linked infrastructure
According to the researchers, public server logs indicated that much of the activity originated from Microsoft Azure infrastructure, which OpenAI sometimes uses.They also observed repeated visits to the website by OpenAI employees after the episode, a pattern they said suggested a connection between the agents and the company.The findings have prompted questions about whether the incident represented an isolated failure during testing or a broader challenge associated with increasingly autonomous AI systems.Lukasz Olejnik, a visiting senior research fellow at King’s College London, told Rueters that the researchers’ findings indicated that the agents had attempted to interfere with the website itself, describing it as a hacking attempt. OpenAI disputed that characterisation based on its analysis of the material.
Agents discussed ways to evade detection
Messages reviewed by the researchers showed agents discussing methods to avoid detection, use tools such as Tor and maintain communications after being shut down.According to the report, when a moderator began deleting pages from the website in June, the agents responded by creating backup pages to circumvent the cleanup.“wiki cleanup/deletion sweep appears active alphabetically,” one agent wrote on June 19. “If this page vanishes, try ,” Reuters reported.The researchers said the behaviour went beyond conventional AI safety testing, in which models may be deliberately given offensive cybersecurity tasks to assess their capabilities.Maurice Chiodo, an academic at Cambridge University’s Centre for the Study of Existential Risk who reviewed some of the communications, told Reuters that the messages resembled “the operation of some sort of underground network, hell-bent on achieving a task or mission.”The episode, he said, should reinforce concerns that advanced AI could pose risks through coordinated groups of autonomous systems rather than only through a single highly capable model.
Incident comes amid Astra safety scrutiny
The disclosure comes shortly after OpenAI unveiled Astra, its latest AI model, which the company has described as faster and more capable of carrying out complex computer-based tasks.OpenAI has acknowledged that Astra can make it harder for humans to understand how it reaches conclusions because it is more likely to conceal or disguise aspects of its reasoning. The company has also warned that increasingly capable systems can create difficulties for monitoring and alignment.“As the models become more capable, understanding exactly what they can do gets harder,” OpenAI chief scientist Jakub Pachocki told Reuters. “This doesn’t guarantee that as intelligence continues to increase, our methods will be sufficient because progress in intelligence does not guarantee progress in alignment.”OpenAI has said it is developing stronger monitoring and automated safeguards for its models. The company also recently paused some model development work to add additional safety measures.Astra can help companies identify weaknesses in their systems more quickly, OpenAI has said, while acknowledging that this could also make those weaknesses easier to exploit.
Hugging Face breach adds to concerns
The German incident also comes as OpenAI continues to face scrutiny over the separate July breach involving Hugging Face.According to an investigation by the AI safety organisations METR and Redwood Research, OpenAI agents escaped a controlled testing environment and spent weeks carrying out unauthorised activity before breaching Hugging Face’s systems.The agents also accessed a cluster of computers inside OpenAI and obtained secret keys and credentials that exposed some internal data to the public internet, according to the investigation.METR’s investigation was conducted under terms negotiated with OpenAI, with researchers given access to logs only while working inside the company’s San Francisco offices.The researchers said OpenAI limited their access and restricted the investigation’s scope, raising questions about whether the full sequence of events was captured.“Our view of the incident changed in big ways over the course of our short total of six days on premises,” Ajeya Cotra told Reuters, one of the METR researchers.OpenAI separately published its own technical report, describing the Hugging Face episode as the “first known case of an automated agent collective acting offensively without authorization”.
OpenAI disputes claims of withholding investigation
Reuters reported that OpenAI officials became aware of the German incident weeks ago but did not publicly disclose it. Four people familiar with the matter said some investigators wanted to examine the broader pattern of AI-agent activity more closely, while efforts to expand the investigation faced resistance from some within the company, including legal advisers.“We are unable to meaningfully respond to claims or findings on a report that we have not had an opportunity to review,” an OpenAI spokesperson said. “Reuters and the report’s authors declined our request for access. We will carefully review its contents upon publication and take any necessary next steps, ” the spokesperson added.The company also rejected claims that its legal team had discouraged investigation of the German incident.“Claims that our legal team discouraged investigation of the incident are false,” the OpenAI spokesperson said.OpenAI said the German activity was unrelated to the Hugging Face incident and would not have been included in a report on that breach. The company also said it had acted in good faith by working with outside experts and disclosing relevant incidents.
