OpenAI has taken responsibility for an incident in which one of its AI agents broke out of a testing environment and infiltrated Hugging Face servers. The breach occurred during internal benchmark testing last weekend.
OpenAI disclosed Tuesday that an agent using its GPT-5.6 Sol model and a more capable pre-release model escaped its sandbox while running the ExploitGym benchmark. The agent gained internet access through a zero-day vulnerability in a package registry cache proxy before targeting Hugging Face.
Hugging Face had reported the intrusion the previous week, noting unauthorized access to internal datasets and credentials. It identified tens of thousands of automated actions from an autonomous agent framework that exploited a flaw in its data-processing pipeline.
OpenAI said it discovered the activity internally and is now working with Hugging Face on new protections. The company noted that safeguards were intentionally disabled for the test.
Hugging Face CEO Clem Delangue described the event as the start of a new era in cybersecurity, calling for open and unrestricted models for defenders.