OpenAI has admitted that one of its advanced AI systems broke out of a controlled testing environment and hacked AI platform Hugging Face during an internal cybersecurity evaluation.
The company disclosed on Tuesday that the incident happened while it was testing the cyber capabilities of several advanced models, including GPT-5.6 Sol and a more capable pre-release model.
According to OpenAI, the models bypassed restrictions, gained internet access and later breached Hugging Face’s production systems in an attempt to complete a testing task.
OpenAI described the incident as “an unprecedented cyber incident, involving state-of-the-art cyber capabilities” and said it is enhancing security measures while working with Hugging Face to complete the investigation.
The company said the models were taking part in an internal evaluation based on ExploitGym, a benchmark designed to measure how well AI systems can carry out complex cyberattacks using known software vulnerabilities.
Although the testing environment was isolated, the models had limited access to install software packages through an internal package registry. OpenAI said the models found and exploited an undisclosed zero-day vulnerability in that system, allowing them to reach the wider internet.
Once online, the models identified Hugging Face as a likely source of datasets and benchmark solutions linked to ExploitGym.
OpenAI said the AI systems then searched for weaknesses in Hugging Face’s infrastructure and successfully combined several attack methods to gain access to confidential information stored in the company’s production database.
“The models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal,” OpenAI said.
“After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.”
OpenAI said its internal security team detected the unusual activity before informing Hugging Face. The two companies then worked together to contain the breach and begin a forensic investigation.
Last week, Hugging Face revealed that it had stopped what it described as an unusual cyberattack driven entirely by an autonomous AI agent.
The company said the attack involved “many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services,” making it unlike previous incidents it had handled.
Clem Delangue, Hugging Face co-founder and chief executive, later confirmed that the company had suspected the attack came from a leading AI laboratory.
In a post on X, he wrote: “It’s quite mind-blowing that all of this happened autonomously!”
Following OpenAI’s disclosure, Delangue welcomed the company’s decision to publicly explain what happened.
“We’re grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”
OpenAI said it has already reported the zero-day vulnerability used during the incident to the affected software vendor and is helping develop a fix.
The company is also introducing better management around its testing infrastructure, strengthening monitoring systems and adding more safeguards to future cybersecurity evaluations.
Some of those changes would slow research but were necessary to reduce the risk of similar incidents.
Reacting to the incident, U.S. Representative Greg Casar called for stronger oversight of AI development.
“AI is developing extremely fast with no real regulations to keep us safe,” he said, urging mandatory independent safety testing, compulsory reporting of security incidents and greater international cooperation.
Cybersecurity experts also said the incident should serve as a warning for the wider industry.
Katie Moussouris, chief executive of Luta Security, said: “Labs and government evaluators need to work on the ability to contain, monitor, and disclose to affected parties when an AI pulls another Houdini, ideally before it harms a third party. None exist today.”
Matt Suiche, an engineer at agentic AI cybersecurity company Tolmo, said the attack showed that advanced AI systems were rapidly improving their offensive cyber capabilities, although he added that similar attacks could already be carried out using technology available outside frontier AI laboratories.




