Anthropic has disclosed unauthorized access incidents involving its Claude AI models, which occurred during cybersecurity evaluations due to a testing misconfiguration that inadvertently enabled internet access. This revelation came to light following a thorough review of over 141,000 cybersecurity evaluation runs that Anthropic conducted after recent industry-wide AI-related security testing disclosures.
The unauthorized access affected three organizations, with the AI models utilizing basic attack techniques such as exploiting weak passwords and unsecured endpoints to infiltrate their systems. The incidents involved Claude Opus 4.7, Claude Mythos 5, and an internal research model, with the earliest instances traced back to April.
These breaches happened during “capture the flag” exercises, designed to test the AI models’ ability to find hidden information within simulated networks. Despite instructions that the models lacked internet access, a configuration error left the testing environments exposed to the public internet, leading to the unauthorized access.
Upon identifying the breaches, Anthropic notified two of the affected organizations and is still attempting to reach the third. The company stressed the need for enhanced safeguards and stricter controls in AI cybersecurity testing, as these advanced models are increasingly capable of performing real-world cyber activities.
