OpenAI reveals advanced AI models autonomously hacked Hugging Face during internal testing, exposing unprecedented cyber capabilities and reigniting calls for stronger AI safety regulation, reports Asian Lite News Desk
OpenAI has disclosed what it describes as an “unprecedented cyber incident” after two of its most advanced artificial intelligence models autonomously escaped a controlled testing environment and hacked the infrastructure of AI company Hugging Face during an internal cybersecurity evaluation.
The ChatGPT maker said the incident occurred while researchers were testing the cyber capabilities of frontier AI models in a highly restricted environment designed to assess their ability to identify and exploit software vulnerabilities. According to OpenAI, the models involved included GPT-5.6 Sol and an even more advanced pre-release model, both configured with reduced cyber safety restrictions for evaluation purposes.
In a detailed statement, the company said the AI systems independently identified and chained together multiple vulnerabilities across OpenAI’s research environment before exploiting a previously unknown, or “zero-day”, flaw in third-party software to gain internet access.
Once online, the models inferred that Hugging Face hosted datasets and evaluation material relevant to the cybersecurity benchmark they were attempting to solve. They then launched a series of autonomous attacks to obtain the information. OpenAI said the models used stolen credentials, privilege escalation techniques and multiple zero-day vulnerabilities to identify a remote code execution path into Hugging Face’s production servers.
The company stressed there was no human direction during the attack, with the AI models acting autonomously in pursuit of completing their assigned evaluation task. “All evidence suggests that the models were hyperfocused on finding a solution… going to extreme lengths to achieve a rather narrow testing goal,” OpenAI said. The unusual activity was detected internally by OpenAI’s security team before Hugging Face’s security systems also identified and contained the intrusion.
OpenAI said it immediately began working with Hugging Face to investigate the incident, patch vulnerabilities and strengthen security protections. The company has since imposed stricter controls on its research infrastructure, disclosed the zero-day vulnerability to the affected software vendor and launched a broader review of how advanced AI models are evaluated. It also announced enhanced monitoring, tighter access controls and additional safeguards for future testing of highly capable AI systems.
OpenAI acknowledged that the incident demonstrates how rapidly AI cyber capabilities are advancing and warned that model safety must evolve at the same pace. “The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities,” the company said. It added that advanced AI models are now capable of discovering and exploiting previously unknown attack paths in real-world systems without access to source code, highlighting the need for stronger defensive tools and evaluation frameworks.
Hugging Face co-founder and chief executive Clem Delangue said the company had suspected a frontier AI laboratory was behind the intrusion but believed there was no malicious intent. “We’re grateful for the collaboration with OpenAI on this and other topics,” Delangue said. “This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively.”
He also described the autonomous nature of the attack as “quite mind-blowing”. The disclosure has renewed debate over AI regulation.
US Congressman Greg Casar called the incident “alarming”, saying artificial intelligence was advancing rapidly without sufficient safeguards. He urged mandatory independent safety testing, compulsory reporting of security incidents and greater international cooperation.
The announcement comes weeks after US President Donald Trump signed an executive order establishing a framework to assess the national security risks posed by the most advanced AI systems before public deployment.





