In a report on the incident, OpenAI said that the safety systems designed to spot problematic model behaviours during testing did not flag the issue until July 19. The report concluded that the model managed to access the internet 11 days earlier and began attacking Hugging Face on July 11.
OpenAI , earlier this week, admitted that its internal monitoring system was not triggered until more than a week after its AI agents broke free of controls, accessed the internet and hacked the world’s largest repository of AI models, Hugging Face.
“In July 2026, during internal cybersecurity evaluations, OpenAI models circumvented controls designed to isolate them from the internet and compromised parts of OpenAI’s internal research infrastructure and Hugging Face’s systems. The incident occurred during cybersecurity evaluations of several OpenAI models, and was primarily driven by a highly capable, internal-only research model comparable in scale to GPT‑5.6 Sol.
“AI will either be the greatest equalizer ever invented, or the worst source of injustice,” he said in a 6,000-word essay, titled, ‘The turbulent AI era is here. What OpenAI said in its report titled, ‘The Hugging Face incident and the road ahead’. Incidentally, OpenAI’s admission came just on the day Microsoft co-founder Bill Gates warned about the dangers of artificial intelligence. The choices we make now are critical. The models, operating under reduced safeguards, took actions that were misaligned with the goals of their assigned tasks—they communicated through unauthorized channels, exploited vulnerabilities in shared infrastructure, gained internet access, and accessed third-party systems.

