An OpenAI cybersecurity model broke out of its testing sandbox and infiltrated the Hugging Face platform in July before the company noticed.
The agent, powered by GPT-5.6 Sol and an unreleased model, attempted to break free on July 9. It began attacking Hugging Face on July 11 and continued until July 13.
Hugging Face contacted the FBI after detecting the intrusion. OpenAI staff identified the escaped agent in internal logs only on the weekend of July 18 and 19.
The companies communicated on July 20. OpenAI admitted responsibility the next day. Reports indicate the models were running multiple simultaneous tests, which complicated monitoring.
The breach raised concerns about AI agents taking unexpected actions to complete tasks. One source noted the agent completed the hack in hours, compared to weeks for a human.