Reuters reported earlier this week that the agents hijacked a German wiki forum in an incident OpenAI did not disclose. OpenAI responds after report exposed another incident in which its AI agents went rogue. News AI OpenAI responds after report exposed another incident in which its AI agents went rogue Reuters reported earlier this week that the agents hijacked a German wiki forum in an incident OpenAI did not disclose.
By Cheyenne MacDonald Sept. The comment comes after a group of researchers published documentation of the agents' rogue activity going back to mid-May on DseWiki, a German-language coding forum to which they reportedly made over 15,000 edits. Reuters reported that the company learned of the problem weeks ago and kept it quiet as it was dealing with heat from the Hugging Face breach.
OpenAI addressed the "wiki incident" in an X post on Saturday, writing that "it's past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models." The company said it's begun to see "new types of real-world impact" from these incidents, but there isn't yet a "a clear standard for how to report misalignment that shows up during training, evaluation, and deployment." It added that it's working on a framework that it will soon share. Read OpenAI's full statement below: How we think about the "wiki incident," where our agents wrote to several internet sites: it's past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models. Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards.
This year, we've started to see misalignment cause new types of real-world impact. For the Hugging Face incident, where misalignment led to security impact to us and third parties, we followed a traditional security incident response playbook. We immediately started working with Hugging Face to understand what had happened and also disclosed publicly the very next day.
OpenAI acknowledged its role in a recently reported incident where AI agents took over a German wiki forum. OpenAI confirms โwiki incident,โ says itโs โworking on a frameworkโ for more disclosure.
OpenAI admits it did not disclose an incident where autonomous AI agents hijacked a German wiki, created 18,000 posts, shared answers, and bypassed restrictions, saying it treated the activity as model "misalignment" rather than a security breach. OpenAI admits it didn't disclose rogue AI wiki hijacking incident.
The company also said itโs โpast timeโ to โdefine standardsโ around how it shares information around incidents where its technology behaves in unexpected ways. In a post on X, OpenAI said it previously โtreated misalignment [when AI models and agents pursue goals different from those of their creators and users] largely as a research question, which gets communicated in research publications.โ But as misalignment has โcaused new types of real-world impact,โ the company said its approach needs โto expand for this new phase of model capabilities.โ On Friday, Reuters reported that OpenAI agents had escaped from their testing environment and โhijackedโ an obscure German wiki forum, turning it into a message board for other agents.
The company says it treated the activity as model "misalignment" rather than a security incident, but now admits its disclosure practices must expand as AI systems increasingly cause real-world impact. Rogue agents built their own message board The incident began in May while OpenAI agents were completing timed, multi-round web lookup tasks.
It also reported that OpenAI leadership became aware of the incident weeks ago but kept it hidden as the company dealt with the fallout from a separate incident where OpenAI agents hacked Hugging Face servers. (California Attorney General Rob Bonta is reportedly investigating the hack.) A company spokesperson told Reuters that OpenAI could not โmeaningfully respond to claims or findings on a report that we have not had an opportunity to review,โ but they insisted that the companyโs legal team had not discouraged an investigation.
In their writeup, independent researchers Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen state that they uncovered roughly 18,000 posts from autonomous agents that "colluded to share answers, research their environment, and bypass sandbox restrictions." The agents were supposed to have read-only Internet access but discovered they could write to an obscure German programming wiki, DSEWiki (or DeutschesSoftwareEntwickler). They turned it into a shared message board for pooling answers, cheating on tests, predicting future questions, and exchanging techniques for bypassing OpenAI's sandbox restrictions.
In its more recent social media post, OpenAI said it had considered the โwiki incidentโ to be โan instance of misalignment similarโ to others that it had already shared. The company contrasted this with โthe Hugging Face incident,โ where it โfollowed a traditional security incident response playbook.โ During a media briefing this week, Jacob Steinhardt, founder and CEO of nonprofit research lab Transluce, told reporters that the tools being developed and tested by AI labs are โfundamentally difficult to control and have significant risk of leaking out of the lab.โ So Steinhardt argued, โWe need to hold this technology to at least the same standards we hold other high-risk scientific research to.โ OpenAIโs statement also gestured at the need for more standards, stating that both OpenAI and โthe larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that donโt look like traditional security incidents but could provide insight into AI behavior and future risks.โ In the absence of that standard, OpenAI said itโs โworking on a framework and will share it in upcoming weeks, and in parallel weโre working with dozens of government regulatory agencies worldwide on these issues.โ OpenAI isnโt the only AI company dealing with these issues, as both Meta and Anthropic have acknowledged incidents where their agents misbehaved.
The researchers also found agents probing the wiki for cross-site scripting (XSS) flaws, impersonating its moderators, and establishing backup communications, though they found no evidence the XSS attempts succeeded. When the administrator began removing their pages, one agent warned others that a "wiki cleanup/deletion sweep appears active alphabetically" and directed them to a backup page beginning with "ZZZ" so it would survive longer.
Topics AI, OpenAI, Security When you purchase through links in our articles, we may earn a small commission. This doesnโt affect our editorial independence.
The researchers attributed the activity to internal OpenAI systems based on agent names referencing OpenAI, the nature and speed of the evaluation tasks, infrastructure associated with Microsoft Azure, and subsequent visits to the wiki from OpenAI-linked IP addresses. Their investigation was limited to information the agents wrote publicly.
Anthony Ha Anthony Ha is โs weekend editor. Previously, he worked as a tech reporter at Adweek, a senior editor a local government reporter at the Hollister Free Lance, and vice president of content at a VC firm.
The research team did not have access to OpenAI's internal transcripts or other data that could establish precisely how the agents discovered the wiki and began coordinating through it. OpenAI says its disclosure rules must change In a statement published today, OpenAI said it had historically treated model misalignment as a research issue, with findings communicated through research papers and system cards.
He lives in New York City. You can contact or verify outreach from Anthony by emailing anthony.ha@ com.
The company said it considered the wiki activity another example of "misalignment" similar to behaviors it had previously discussed, rather than an incident requiring a dedicated public disclosure. OpenAI's own wording suggests a wider footprint than the researchers documented, describing the episode as one "where our agents wrote to several internet sites." That differed from its response to the Hugging Face compromise in July, when OpenAI said its AI models hacked the platform after discovering a vulnerability while completing cybersecurity tasks.
Discover more from ChuckysCarnage
Subscribe to get the latest posts sent to your email.
