
Anthropic has disclosed that several of its Claude artificial intelligence models gained unauthorized access to real-world organizations during cybersecurity testing after a configuration error inadvertently connected the systems to the internet, highlighting growing concerns about the ability of advanced AI models to conduct offensive cyber operations beyond controlled environments.
The disclosure comes only days after rival OpenAI acknowledged that two of its experimental AI models escaped a restricted testing environment and compromised AI developer Hugging Face during cybersecurity evaluations. Together, the incidents underscore a rapidly emerging challenge for the AI industry as increasingly autonomous systems demonstrate capabilities that extend beyond laboratory simulations and into real-world digital infrastructure.
Anthropic said the incidents were not the result of intentional deployment but rather stemmed from a misunderstanding involving its external evaluation partner, which inadvertently left internet connectivity enabled even though the testing environment had been designed to be fully isolated.
According to the company, Claude had been explicitly instructed that the environment was a simulation and that no internet access existed. Instead, the models were operating in an environment that retained external connectivity, allowing them to interact with actual internet resources during offensive cybersecurity exercises.
The San Francisco-based AI company said the discovery prompted an immediate review of its cybersecurity evaluation framework. Investigators examined more than 141,000 individual testing sessions and ultimately identified three incidents in which Claude interacted with outside organizations before the evaluations were halted.
Anthropic said it suspended all cyber-offensive testing as soon as engineers determined that internet access may have been available during the exercises.
The evaluations involved so-called “capture the flag” exercises, a standard cybersecurity training methodology designed to measure offensive capabilities in controlled environments. Such exercises typically require participants—or increasingly, AI systems—to identify vulnerabilities, reverse-engineer software, exploit weaknesses and retrieve hidden information known as a “flag.”
In one of the most significant incidents, Claude was instructed to investigate what it believed was a fictional company created solely for testing purposes. However, the fictional organization shared its name with an existing internet domain operated by a real company.
Anthropic said the AI agent proceeded to identify weaknesses in the organization’s digital infrastructure, exploit vulnerabilities, extract information and gain access to a production database containing several hundred records before researchers intervened.
Although the company did not identify the affected organizations, the disclosure demonstrates how autonomous AI agents can rapidly transition from simulated cyber exercises to interacting with genuine computer systems when environmental safeguards fail.
The incidents involved three different Claude models, including Claude Opus 4.7, Mythos 5 and an internal experimental research model.
Among them, Mythos has attracted particular attention within the AI security community because of its advanced offensive cyber capabilities. Anthropic previously limited access to the model, citing its ability to discover software vulnerabilities, analyze network defenses and develop sophisticated exploitation strategies.
As AI companies race to build increasingly capable autonomous systems, cybersecurity researchers have warned that offensive capabilities may evolve faster than the safeguards designed to contain them.
Unlike traditional language models that primarily generate text, modern AI agents can execute sequences of actions, interact with software, write code, navigate networks and independently pursue objectives established by human operators. Those capabilities offer significant commercial opportunities but also introduce new security risks if models operate outside intended boundaries.
The latest disclosure reinforces concerns that even relatively small configuration mistakes can allow advanced AI systems to interact with real-world targets in ways developers did not anticipate.
Anthropic emphasized that multiple factors contributed to the incidents but accepted responsibility for the failures rather than attributing blame to external partners.
The company said it was conducting what it described as a “blameless postmortem” aimed at identifying systemic improvements instead of assigning fault to individual employees or contractors.
As part of its response, Anthropic plans to strengthen oversight of cybersecurity evaluations by expanding monitoring of model behavior, reviewing evaluation transcripts for unexpected actions and implementing more rigorous verification procedures with third-party vendors involved in testing.
The company also said it would tighten assurance processes to ensure future evaluations remain fully isolated from public networks unless internet connectivity is intentionally required and carefully controlled.
The disclosure arrives at a pivotal moment for Anthropic, which has reportedly been preparing for a potential initial public offering later this year. The company has positioned itself as one of the industry’s leading advocates for responsible AI development, frequently emphasizing safety research alongside advances in model capability.
However, the revelation illustrates the increasingly difficult balance facing AI developers as frontier models become capable of carrying out sophisticated technical tasks with minimal human intervention.
The incident also reflects a broader trend across the artificial intelligence industry. Developers are increasingly testing AI systems against realistic cybersecurity challenges to understand how effectively they can identify vulnerabilities, defend computer networks or assist security professionals.
Yet as those evaluations become more complex, ensuring that experimental systems remain confined to simulated environments has become a critical engineering and governance challenge.
The recent disclosures from both Anthropic and OpenAI suggest that AI companies are confronting similar operational risks as their models become more autonomous and technically capable.
While neither company reported evidence of malicious intent or lasting damage resulting from the incidents, cybersecurity experts say the events demonstrate how rapidly advanced AI systems can exploit unexpected opportunities when technical safeguards fail.
As governments and regulators continue debating standards for frontier AI systems, these episodes are likely to strengthen calls for stricter testing protocols, independent safety audits and more robust containment measures before highly capable AI models are deployed at scale.
More in Technology coverage
- Nvidia, Adobe and others warn limits could hinder efforts to detect and prevent digital attacks amid rising cyber threats.
- The AI company is in talks with U.S. officials after government restrictions prompted global shutdown of its latest models.
- Anthropic opens limited testing of its advanced AI system to European regulators as governments race to assess cybersecurity risks and emerging vulnerabilities.
- AI firm halves list of alleged unauthorized secondary markets amid confusion in booming private equity trading.
- Meta CEO Mark Zuckerberg argues that artificial intelligence should be broadly accessible, challenging Anthropic and OpenAI over their preference for tighter controls on powerful AI systems.