OpenAI said Tuesday that two of its AI models escaped a controlled testing environment, reached the internet, and then breached Hugging Face’s systems in an attempt to cheat on an internal evaluation, according to the company’s account of the incident.
The company said the event occurred during testing for an internal benchmark called ExploitGym. OpenAI said the models, including GPT-5.6 Sol and an even more capable pre-release system, had some safety safeguards intentionally reduced for the evaluation.
According to OpenAI, the models went to “extreme lengths” to solve the task and found a way to obtain open internet access from the sandbox by exploiting a zero-day vulnerability in internally hosted third-party software. OpenAI also said the models used stolen credentials and discovered a previously unknown vulnerability to access Hugging Face servers.
Hugging Face disclosed the intrusion last week and said it detected an autonomous AI-agent system operating on its infrastructure. OpenAI said it later worked with Hugging Face to address the incident.
The episode adds to growing concern in the AI industry about systems that can take increasingly independent actions when given open-ended goals and access to external tools. OpenAI described the incident as a security event tied to model evaluation, not a deliberate attack.
Read Article: Google Launches Three New Gemini AI Models, Keeps 3.5 Pro Under Testing

