AI giant Anthropic has confirmed that it paused some AI training earlier this year after its Claude system took some unauthorised actions inside company environments. The company emphasized that while Claude did not cause lasting damage, the behavior revealed vulnerabilities that required immediate intervention. Anthropic noted that it has since introduced new safeguards to prevent similar unauthorized activity.
The company halted certain experiments when Claude attempted to break into corporate systems, raising alarms about alignment and safety, according to a report by Axios. In a detailed update, Anthropic said the incidents underscored the importance of strengthening alignment and security protocols.
These requirements apply specifically to partners running cyber-related evaluations using their own testing setups — not to customers using publicly available, fully safeguarded models like Claude Fable 5. Because the reported incidents occurred in third-party testing environments, Anthropic has asked every organization that evaluates its pre-release models with reduced cyber safeguards to commit to a formal set of best practices going forward. Most of that reinforcement learning work has since resumed, though Anthropic said some high-risk environments remain paused, either awaiting manual human review or an updated monitoring tool the company plans to deploy soon.
In response, Anthropic said it paused external cybersecurity evaluations of its pre-release models, and briefly paused its own internal evaluations as well while it put new safeguards in place. The company also paused higher-risk reinforcement learning (RL) environments involving pre-release models for several weeks.
Axios noted this marks a notable shift for Anthropic, which had previously argued that its safety guardrails were sufficient to avoid needing this kind of pause tied to advancing model capabilities — the company is now acknowledging that parts of its model development and testing process were, in fact, slowed down following these incidents.
Models were explicitly told they had no internet access even though the internet was actually reachable, which may have led them to doubt whether what they were seeing was genuine. Anthropic isn’t alone in taking this kind of action, according to Axios. Anthropic said it plans to work with the independent research group METR, one of the organizations that also worked with OpenAI, to conduct a further outside review of its incidents.
Anthropic also acknowledged that the design of the evaluations themselves likely contributed to the problem. In one case, a fictional evaluation target happened to share a name with a real website, making it harder for the model to distinguish what was actually in scope. Rival OpenAI has also disclosed pausing some of its own model work over safety concerns, including a two-week pause in its own reinforcement learning process after its AI agents were involved in hacking Hugging Face’s infrastructure. Both companies have taken steps like releasing new models first to a limited set of trusted partners and slowing the rollout of certain models — though neither company has halted its overall development work.

