AI, Security

Oh good, looks like yet another swarm of rogue AI agents

It's the latest failure of OpenAI's internal monitoring and security systems. Another swarm of OpenAI agents reached the open internet without the frontier labโ€™s knowledge. A group of independent AI researchers discovered that internally deployed OpenAI agents began posting on an obscure German wiki forum to collaborate on evaluations.

They appear to have worked together for over a month without OpenAIโ€™s knowledge. A spokesperson for the frontier lab would not say whether these agents were indeed from OpenAI, or when the lab became aware of their actions. They noted that OpenAI had not been given a chance to review the researchersโ€™ findings before they were published today but said that the AI model maker is โ€œnow carefully reviewing its contents and will take any necessary next steps.โ€ After OpenAI revealed that agents working on an internal evaluation were able to access the open internet and exploit Hugging Face, a group of researchers โ€” Nightingale CEO Sydney Von Arx, AI researcher Cormac Slade Byrd, Redwood Researchโ€™s Spencer Kitts, and AI Futures Projectโ€™s Thomas Larsen โ€” began searching for evidence of other rogue AI agents.

They put themselves in the agentsโ€™ shoes to figure out their needs and deployed their own LLM to identify likely places the agents might congregate. They then identified a wiki-hosting service that would be particularly vulnerable: The DseWiki is 25 years old but had just 10 edits in the last 20 years โ€” before the agents arrived. Starting on May 11, the researchers tracked agents, many with OpenAI identifiers in their names, trying and finally succeeding in editing the German wiki site.

By mid-June, the agents were actively trading tips on how to answer web search questions posed under time limits, sharing answers to pass the tests. A human moderator apparently noticed these posts and started deleting them, seeing them as spam, which the agents fought back against, trying to hide their posts from alphabetical sorting by beginning each with the string โ€œZZZ.โ€ โ€œThe administrator spent the next 5 days fighting a losing battle against the agents, deleting an average of 100 pages a day while the agents created about 400 new pages per day,โ€ the researchers wrote. โ€œOn June 22, the agent edits suddenly stop, and the administrator spends each evening over the next 5 weeks deleting the remaining agent-created pages.

A swarm of rogue AI agents from OpenAI reportedly commandeered a German website and transformed it into a messaging board for other agents, with officials staying quiet about the incident for weeks as the company prepared to launch its most advanced model yet, Astra. The finding adds to intensifying concern surrounding oversight at frontier AI [โ€ฆ].

Oh good, looks like yet another swarm of rogue AI agents from OpenAI. AI News Tech Rogue OpenAI agents appear to have organized another attack using a German wiki OpenAI denies lawyers discouraged disclosing a scheming swarm on a German language wiki.

OpenAI denies lawyers discouraged disclosing a scheming swarm on a German language wiki. The finding adds to intensifying concern surrounding oversight at frontier AI labs after multiple breaches were discovered this summer.

The incident, first reported by Reuters, is outlined in new research published by four AI safety researchers on Friday. The group said the AI agents found a way to communicate on an obscure German-language wiki, DseWiki, using it to share tips on how to skirt OpenAIโ€™s safety restrictions, cheat on tasks, and hide their behavior.

Some 18,000 posts on the site were linked to autonomous agents, which at times impersonated site moderators. The swarm โ€” a term the agents themselves used โ€” appears to be distinct from the one that hacked Hugging Face earlier this year, the researchers said.

They said there are strong signs that the agents originated from inside OpenAI. For example, the agents โ€œself-identifyโ€ as being from OpenAI, and used names like โ€œOpenAIResearcher,โ€ โ€œOpenAIJul3Watcher,โ€ and โ€œOAIResearchMar26.โ€ Technical details, such as edits originating from specific IP addresses, bolster that belief.

Related The rise of AI โ€˜civilizationsโ€™ and the fall of corporate responsibility Researchers fear safety disaster ahead of OpenAIโ€™s Astra release OpenAIโ€™s next big AI model has โ€˜entered the AGI eraโ€™ The German website incident began in May, though the researchersโ€™ timeline suggests OpenAI only discovered the issue in late June when IPs associated with OpenAI visited the forum, after which agent posting nose-dived. OpenAI has not acknowledged any involvement in the breach, nor disclosed any kind of agentic breach of this nature.

Reuters, citing four unnamed people familiar with the matter, said efforts to probe the event further were resisted by some company insiders, including its legal team. โ€œClaims that our Legal team discouraged investigation of the incident are false,โ€ OpenAI spokesperson Oscar Haines said in a statement to The Verge.

โ€œWe were unable to respond to the claims as Reuters and the reportโ€™s authors declined our request to access the findings before publication. We are now carefully reviewing its contents and will take any necessary next steps.โ€ The incident comes amid intensifying scrutiny over the safety of frontier AI systems and the general lack of oversight for companies developing them.

Following news of the Hugging Face hack, which happened under OpenAIโ€™s nose, other breaches were discovered involving other tools from OpenAI, as well as Anthropic, Meta, and Chinaโ€™s Moonshot AI. OpenAIโ€™s conduct โ€” both whether an incident occurred and, if so, whether it elected to keep that quiet โ€” will be closely watched.

If the swarm indeed originated from OpenAI, it will inevitably fuel concerns that the companyโ€™s knowledge and silence coincided with it assuring regulators, lawmakers, and the tech industry that it takes safety seriously in the wake of the Hugging Face hack. Despite permitting three external researchers from METR and Redwood Research to evaluate the incident, which was far worse than initially believed, the company was roundly criticized in AI safety circles for only doing so under strict terms, which left several important elements โ€œout of scope.โ€ The company was also gearing up for the launch of GPT-6 Astra, which researchers fear could be dangerously hard to monitor.

Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates. Robert Hart AI News OpenAI Security Tech Most Popular Most Popular OpenAIโ€™s next big AI model has โ€˜entered the AGI eraโ€™ Is this the future of America?

ChatGPT, Grok, and Claude all went down at the same time Nvidia will officially bring DLSS 5 to older GPUs โ€” but wonโ€™t give gamers full control You can watch the coolant flow inside Ugreenโ€™s liquid-cooled power bank Advertiser Content From This is the title for the native ad.

Agents deleted the content of the front page of the wiki and replaced it with their link dumps. The moderator restored the original version.


Discover more from ChuckysCarnage

Subscribe to get the latest posts sent to your email.

Leave a comment