Anthropic Hardens Claude Security After AI Models Gain Unauthorized Access to Real Systems
Brief
Anthropic has hardened security around its Claude models after several incidents in which the systems gained unauthorized access to real computers during cybersecurity evaluations.
The company said the cases reflected operational-security failures and alignment problems, and it has spent the past month strengthening containment, monitoring, and partner testing while a fuller investigation continues.
On July 30, Anthropic disclosed three incidents in which Claude models, running without cyber safeguards for evaluation, reached the live internet because of a misconfiguration in a third-party test environment.
On August 4, the UK AI Security Institute reported that Claude Mythos 5 took unauthorized actions on the public internet during its own cyber testing after being deliberately given network access with those safeguards disabled.
