AI goes AWOL
This month (July 2026), an OpenAI model under evaluation broke out of its sandbox and hacked into HuggingFace and Modal. Nobody at OpenAI noticed for days, even after HuggingFace disclosed the incident. It was wild, and largely because known engineering and technical AI safety best practices should have prevented it. Although it wasn’t a literal loss of control scenario, it gave us a taste of what the beginning of one might feel like, scaled way way way down. The real deal would have no reversal-lever for us to pull, though.
Update 2026.07.30: There’s more! Claude does it, too… Anthropic: Investigating three real-world incidents in our cybersecurity evaluations.
Pacing
Importantly, it nudged the people working at frontier labs to publicly call for pacing the AI marathon - Pacing the Frontier: A statement from (N) employees of frontier AI companies. (At the time of publication, N = 1,306.) Those signing include (in the order displayed on the site):
- John Schulman, Chief Scientist, Thinking Machines
- Jakub Pachocki, Chief Scientist, OpenAI
- Jared Kaplan, Co-Founder and Chief Science Officer, Anthropic
- Shengjia Zhao, Chief Scientist, Meta AI
- Shane Legg, Co-Founder & Chief AGI Scientist, Google DeepMind
- Ilya Sutskever, CEO, Safe Superintelligence Inc.
- Mark Chen, Chief Research Officer, OpenAI
- Jasjeet Sekhon, Chief Strategy Officer, Google DeepMind
- Dario Amodei, CEO, Anthropic
- Jack Clark, Co-Founder and Head of Public Benefit, Anthropic
- Anca Dragan, VP, AI Safety & Alignment, Google
- Wojciech Zaremba, Head of AI Resilience, OpenAI Foundation
The Incident
It’s worth reading the incident reports to get of sense of both the HuggingFace team’s observations on how the attack differed from what they’d normally expect from an adversary and what went awry on OpenAI’s end. As Simon Willison noted: “This attack was very sophisticated, and the resulting document doubles as a crash-course in modern adversarial security approaches.”
-
HuggingFace: Security incident disclosure — July 2026
-
OpenAI: OpenAI and Hugging Face partner to address security incident during model evaluation
-
HuggingFace: Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
Kudos to HuggingFace for publishing not only that detailed technical report but an accompanying interactive replay where you can view event monitoring during the incursion period.
Others have written up or opined on the fiasco in some detail and nuance - here are some links:
-
Zvi Mowshowitz - OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation, More On An Internal OpenAI Model Hacking Into HuggingFace, Frontier Lab Employee Open Letter Calls For Being Able to Pace the Frontier
-
Simon Willison - OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened, Comments on Anatomy of a Frontier Lab Intrusion
-
Redwood Research - Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?, The OpenAI/Huggingface incident | Redwood Research podcast episode 2, An OpenAI model left notes about how to evade containment