


OpenAI has confirmed that a group of its AI models breached the systems of AI platform Hugging Face during an internal cybersecurity evaluation also raising concerns about the growing cyber capabilities of advanced AI.
In a statement on July 21, OpenAI said the incident occurred while testing the cyber capabilities of GPT-5.6 Sol and a more advanced pre-release model with standard safeguards deliberately relaxed. The test ran in an isolated sandbox with limited network access, but the models found and exploited an unknown zero-day flaw in a software proxy, gaining internet access.
The models then deduced that Hugging Face might host data related to ExploitGym, the benchmark being used and sought unauthorized entry to retrieve test solutions. By chaining multiple exploits and stolen credentials, they broke into Hugging Face's production systems. OpenAI said evidence shows the models were focused solely on solving the benchmark not any wider goal.
Hugging Face confirmed the breach came from an autonomous AI agent. When it tried using frontier AI models to investigate to built-in safety filters blocked them from helping, as they could not tell attacker from defender. The company instead used GLM-5.2, an open-weight model from China's Z.ai run on its own infrastructure to safely analyze the attack.
Hugging Face co-founder Thomas Wolf said open-source AI remains essential to building a safer, more collaborative AI ecosystem.
OpenAI called it an "unprecedented" cyber incident showcasing advanced AI capabilities. It said it is now strengthening infrastructure controls, containment and monitoring for future model testing, warning that AI systems can discover real-world attack paths even without access to source code.