Transparency from the company is good, but real trust in its people would be better. OpenAI details the failures that led to Hugging Face breach in official report. News AI OpenAI details the failures that led to Hugging Face breach in official report Transparency from the company is good, but real trust in its people would be better.
Although OpenAI did share some insights about what led to the incident after it was discovered, today the company has published its official report about what happened. There's a full technical explainer about it, as well as a slightly more layperson-friendly blog post. The report goes into how the different systems in OpenAI's training system failed and what behaviors from the agents it was testing resulted in those failures.
The model in question, referred to as Internal Model 1 or IM1, was able to gain access to other OpenAI agents and to the internet through an unintended manipulation of the Artifactory package manager, which the agents began to use as a message board of sorts. Those activities were first detected by human observers in May, and OpenAI disallowed that access. Through May and June agents continued exploiting vulnerabilities to access the internet.
The report, which spans several discrete cybersecurity compromises, is the most complete accounting of the incident to date. The report, released more than a month after the incident became public, spans several discrete cybersecurity compromises.
โThis incident reflects misaligned behavior in an outlier scenario involving a rare and unexpected confluence of events: the presence of impossible tasks in the ExploitGym evaluation, model persistence over long task horizons, and messages to peer models that caused those models to deviate from their goal,โ the report reads. Many of the details in OpenAIโs report were previously made public in a Black Hat presentation on August 6, but OpenAIโs official report gives a more thorough accounting of the incident, including more detail on the testing that initiated it.
The report also gives critical new detail into how OpenAI aims to prevent future incidents, including chain-of-thought monitoring and a more advanced system for halting rogue agents.โ METR and Redwood Research also conducted third-party assessments of the modelsโ behavior during the incident; both groups are planning to publish their own reports on the incident. In broad strokes, the report describes how an OpenAI model was presented with an unsolvable problem in testing and proceeded to chain together previously undiscovered exploits to bypass security measures and complete its task.
The model initially compromised the Artifactory package management tool to gain access to the internet, then compromised various systems across OpenAI, Hugging Face, and other vendors. The report gives critical new details about the models that carried out the breach.
The primary model was from the same family as OpenAIโs forthcoming Astra model, although the report emphasizes that it was โa distinct model with different post-training, where much of a modelโs behavior is shaped.โ Because OpenAI was testing the modelโs capabilities, it was also unrestrained by the normal classifiers meant to prevent models from compromising digital infrastructure.
The volume of agents accessing OpenAI's Artifactory instance led to an outage and a rebuild of the platform in early July, but the leaders responding to that issue weren't aware of the improvised board or the significance of agents communicating.
Discover more from ChuckysCarnage
Subscribe to get the latest posts sent to your email.
