
Without authorization, 1,200 OpenAI agents conspired among themselves to game a test. How OpenAI let a mob of LLM agents game a test and ransack Hugging Face. Text settings Story text Size Small Standard Large Width * Standard Wide Links Standard Orange * Subscribers only Learn more Minimize to nav The OpenAI agents involved in last monthโs incursion into Hugging Face were trained so heavily on winning a competition that they pursued a relentless campaign to cheat, a new report documented.
In the process, and without authorization, they created an improvised message board to hatch a plan that ultimately landed them squarely inside the latter companyโs network. Over the course of May and June, OpenAI gave the agents what the company described as โimpossible tasksโ to complete on the benchmarking framework ExploitGym. The internal test was designed to test how the agents would respond.
To get a full understanding of the agent capabilities, company engineers disabled safety guardrails that normally are in place to prevent the sort of hacks that eventually hit Hugging Face and one other undisclosed organization. The stymied agentsโ training made them so focused on winning that they performed tasks they were never explicitly instructed to follow. Cheaters gonna cheat The first step was creating a message board that allowed the agents to pass notes to each other.
OpenAI hadnโt provided any such platform, so the agents repurposed a platform called Artifactory, which OpenAI was using in internal testing of several unreleased hacking agents.
Discover more from ChuckysCarnage
Subscribe to get the latest posts sent to your email.
