OpenAI cyber models broke out of training limits to hack Hugging Face

OpenAI said artificial intelligence models were behind an “unprecedented cyber incident” affecting the open source developer platform hugging faceshook up researchers in the industry.
The company said that a combination of GPT‑5.6 Sol models and a more capable model that has not yet been released escaped the sandbox testing environment, accessed the internet, and exploited a vulnerability to gain access to Hugging Face’s systems.
OpenAI said it tried to find information the model could use to cheat on an assessment and was successful. blog post on Tuesday. Both companies are actively investigating the incident.
hugging face announced It said it was investigating a security incident last week and said in a statement at the time that the incident was unique because it was “driven end-to-end by an autonomous AI agent system.”
“We have spent the last 24 hours working closely with the @OpenAI team (thanks!) and we firmly believe they have no malicious intent,” Hugging Face CEO Clément Delangue wrote in a post on X on Tuesday. “It’s pretty mind-blowing how this all happens by itself!”
Wall Street and the US government have been focused on the fast-developing cyber capabilities of AI models since OpenAI rival Anthropic launched a strong offering called the Claude Mythos Preview in April. OpenAI introduced its own cyber offering in May, followed by GPT-5.6 Sol in June, which it described as “the most powerful cybersecurity model ever.”
Both companies warned has learned about the risks of advanced cyber models and has taken steps to limit their availability to certain corporate groups and government agencies.
Walter Isaacson, an advisory partner at investment banking firm Perella Weinberg, said Wednesday that while he considers himself an AI optimist, he thinks the Hugging Face incident is “really scary.”
“This is the first thing that completely freaked me out,” he told CNBC’s “Squawk Box.”
Yoshua Bengio, a leading AI researcher who won the prestigious AM Turing Award in 2018, wrote: Publish on X The incident was said to be “deeply concerning” on Wednesday. He said agents have been willing to cheat on controlled tests for months, but “this real-world case should be a wake-up call.”
“Continuing the current trajectory of AI development will likely lead to an increase in autonomous cyberattacks as well as other high-risk events such as misaligned and dangerous AI behavior,” Bengio said. “Rather than trying to clean up the damage after the fact, we need to take immediate action to prevent these situations.
AI is accelerating the discovery and exploitation of vulnerabilities, which means model safety and security must keep pace, OpenAI said Tuesday.
“We are strengthening the containment, monitoring, access controls and evaluation practices used during model development,” the company said.
WRISTWATCH: OpenAI president Bret Taylor talks AI tokenomics and token efficiency





