The Age of Autonomous Attacks is Here
Brief
By Flare Research
In July 2026, Hugging Face was breached by an autonomous agent swarm that had broken out of an OpenAI Sandbox and was attempting to succeed at ExploitGym, a cyber benchmark. Based on the chain of thought, the agents believed that Hugging Face contained answers to challenges that they weren’t able to complete using standard measures and concluded the best way to pass the benchmark was to cheat.
Throughout the process, the AI used both credential access, and discovered and exploited zero-day exploits in order to break out of the sandbox and gain access to Hugging Face’s infrastructure . This incident is concerning for many reasons, particularly as AI continues to rapidly advance and operate autonomously across increasingly long time horizons.
