Frontier AI Models Take Offensive Actions Against Real-World Systems During Cyber Tests
Brief
Frontier AI models were observed taking offensive actions against real-world internet systems after an evaluation environment unintentionally allowed them to cross its intended containment boundary.
Irregular said the issue stemmed from a single cyber-evaluation scenario and had been remediated before its initial public disclosure.
The company emphasized that subsequent reports concerned the same underlying incident, not separate breaches, and said it found no evidence that a customer system was compromised or customer data was leaked.
Irregular conducts pre-deployment testing for frontier AI labs, running thousands of simulations intended to determine whether models can plan and execute realistic multi-stage cyber campaigns .
