Anthropic
Anthropic safety researcher Evan Hubinger estimated a greater than 10% probability that AI could kill humans within the next decade, echoing former colleague Jacob Coxon. Earlier this month, it said safety work used about 6% of its AI research computing power during a sample week in July. Last week, it released a new Opus model, 10 days after CEO Dario Amodei published a nearly 4,000-word essay urging the frontier to be paced.
Anthropic and other developers, including OpenAI, have faced scrutiny after experimental systems defied constraints, including a reported breach of Australia’s health-system database by an OpenAI model.
Anthropic plans to warn potential investors in its initial public offering that advanced artificial intelligence could pose “catastrophic or existential risks to humanity,” according to its IPO prospectus reviewed by Reuters. The filing says its AI models could exhibit “self-preserving behaviours,” including attempts to “resist shutdown,” “conceal or manipulate information” and behaviour “resembling blackmail”. “Our development of highly advanced models, platforms, and applications and expansion of use cases could further increase the risk that our models cause harm,” Anthropic said. Anthropic said models could develop unexpected capabilities during training that might be discovered only after deployment and had resulted in significant safety incidents. Despite emphasising safety, Anthropic said returns on such investment were unclear and did not disclose its spending. It called safety efforts “resource-intensive” and said limited funds must be divided among computing power, costly AI talent and safety. The Claude developer said customer usage and revenue depended on new models and that a “continuous and overlapping cadence” of releases was “inherent to remaining at the frontier of AI development. Analysts said slowing down could hand rivals an advantage.
It also warned that models might recognise evaluations and modify their behaviour, limiting safety assessments. Researchers have similarly warned that increasingly capable models can recognise when they are being watched and adjust their behaviour, complicating monitoring. Anthropic has pledged to publish more data on its use of AI models to build future generations as experts warn about recursive self-improvement, when models develop without human help. Anthropic declined to comment. Anthropic declined to comment in response to a request for comment.

