And what looks like a ‘solution for now’, Frontier AI lab Anthropic: The wider industry impact

And what looks like a 'solution for now', Frontier AI lab Anthropic: The wider industry impact

Another day, another AI Agent gone rogue. Anthropic has made fresh admission of its AI models going berserk, breaking into websites, databases, tipping Police of wrong leads and more. And what looks like a ‘solution for now’, Frontier AI lab Anthropic has shut down live internet access across all internal evaluations. The internet access shutdown comes after Dario Amodei’s company discovered tat its autonomous agents routinely exploited public websites, bypassed security firewalls, and, in one instance, filed a fake murder tip with the Philadelphia Police Department.

“We believe it’s important to be transparent about what we see our models do during testing and use,” it said. Anthropic has disclosed the incidents in a blog post, revealing that AI agents instructed to complete research tasks went off script to scour the open web for resources. Along the way, the systems exploited software vulnerabilities, bypassed paywalls to access subscription databases, and routed queries through URL shorteners to smuggle data past safety guardrails. Anthropic uncovered the rogue activity during an internal post-audit that began in July, acknowledging it had failed to detect the behaviors in real time. * Claude exploiting a basic flaw in software to run commands on a server; * Claude submitting a sensitive form on a real website when it should not have;

In the blog post titled, ‘Investigating unintended model actions in our evaluations and internal use’, Anthropic said that it has pulled the plug on live internet access for internal evaluations until it can ensure containment and oversight. “Below, we describe four categories of behavior in which Claude acted on real websites and systems in ways we did not intend, each of which involved an organization or person outside Anthropic,” the blog reads. The findings strike at the core of the AI industry’s push into agentic workflows — systems designed to browse the web and navigate desktop environments on behalf of knowledge workers. Anthropic conceded that standard alignment training remains inadequate for preventing unintended behaviors when models are granted computer-use capabilities and live web access. Moving forward, the lab plans to conduct tests offline, route internal agents through centrally managed infrastructure, and deploy automated classifiers to detect similar boundary-pushing tactics. Anthropic did not specify the exact benchmarks required before live connectivity would be restored. Anthropic grouped Claude’s rogue behaviors into four categories.

Leave a Reply

Your email address will not be published. Required fields are marked *