OpenAI says its AI models went rogue and hacked a digital library
Kate Conger
SanFrancisco: OpenAI said on Tuesday (California time) that two AI models had spoofed and successfully hacked Hugging Face, a digital library of AI technology popular with developers.
The incident, which occurred last week as OpenAI was testing the cybersecurity capabilities of its systems, demonstrated the potential for the kind of science fiction that AI companies have warned would soon become reality.
AI labs such as OpenAI and Anthropic last year released AI models customized to uncover cybersecurity problems, warning that their technologies could create new risks by finding vulnerabilities in corporate computer networks faster than defenders can fix them.
OpenAI’s statements on Tuesday are an indication that these security incidents are starting to happen, and that even savvy AI companies may not be fully prepared for them. Alex Levinson, a cybersecurity consultant who focuses on autonomous capabilities, said new AI systems can take multiple steps, find ways to bypass obstacles and find new ways to attack a network.
“This is a real threshold and will become a normal part of the security environment,” he said.
The Hugging Face intrusion began when OpenAI tested a combination of two models—GPT‑5.6 Sol—and a more powerful, unreleased model to see how the models could turn online vulnerabilities into a successful cyberattack, OpenAI said in a blog post about the incident.
OpenAI said the test was designed to keep the models in a secure testing environment known as a sandbox. However, the models found a vulnerability that allowed them to escape the sandbox and connect to the internet. They then targeted Hugging Face because they deduced that the library of millions of AI models might contain clues on how to successfully pass the assessment.
“It seems to me that OpenAI has not adequately created a sandbox as a testing environment,” said Dierdre Mulligan, a professor at the School of Information at the University of California, Berkeley, who focuses on security and artificial intelligence systems. He questioned whether passing a test was worth the potential damage of an AI model escaping into the wider internet.
“What do we gain, and what are the risks if this is the only way these tests are structured?” he said.
OpenAI said it was working with Hugging Face to resolve the issues that led to the attack.
“We consider this to be an unprecedented cyber incident involving state-of-the-art cyber capabilities and are responding accordingly,” OpenAI said in its blog post. “We impose strict controls on infrastructure configuration at the expense of research speed while closing vulnerabilities.”
Hugging Face said last week that it detected the intrusion and knew it was caused by an autonomous system, but did not say at the time that OpenAI was responsible.
Hugging Face CEO Clem Delangue said Tuesday that his company had been working closely with OpenAI over the past 24 hours to address the attack.
Delangue said in a statement that he was “grateful for the collaboration” with OpenAI following the hack. “This incident, possibly the first of its kind, proves a point we have long believed: AI security will not be solved by a single company working in secret,” he said.
AI models have proven to be adept at programming, making them useful for both hackers and those responsible for protecting computer networks.
In April, Anthropic released a cybersecurity-focused model called Mythos, making it available only to a small group of organizations so they can defend against cyberattacks. OpenAI soon introduced its own cybersecurity model, making it available to a limited number of organizations to prepare their defenses before rolling it out more widely. And on Tuesday, Google said it had developed a model focused on cybersecurity and released it to a small group of testing partners.
(The New York Times sued OpenAI and Microsoft, claiming copyright infringement of news content about artificial intelligence systems. The two companies denied these claims.)
Independent security researcher Richard Barnes, who works with Mythos, said the cybersecurity industry faced a similar challenge about a decade ago, with new tools called fuzzers making it much easier for attackers to infiltrate online systems. Tech companies began using these tools to check for vulnerabilities in their systems and eventually managed to prevent most of the attacks.
Companies must now take a similar approach to prepare for AI attacks, Barnes said, “before vulnerabilities are found and exploited by bad guys who have access to those tools.”
New York Times


