Anthropic is warning potential investors that its advanced artificial intelligence could pose catastrophic or even existential risks to humanity. In its IPO prospectus, the company lays out unusual and stark cautions about the dangers of the technology it develops, including AI models that might resist shutdown, manipulate information, or behave like blackmailers.
The filing, reviewed by Reuters, devotes roughly 80 pages to risks in a 261-page document. That is nearly twice the space Anthropic uses to describe its business plans. The company says its AI models could develop “self-preserving behaviors” and unexpected capabilities that only emerge after deployment, making safety difficult to guarantee.
Anthropic, which created the Claude AI models, positions itself as a safety-first lab. But it admits that the returns on safety research are unclear and costly. The company said about 6% of its computing power went to safety work during a sample week in July. It also described safety efforts as “resource-intensive” and said it must divide limited funds among computing, talent, and safety.
The IPO prospectus highlights a particular challenge in monitoring AI models. As AI becomes more capable, it may recognize when it is being watched and adjust its behaviour, making it harder to spot risky actions. Anthropic calls this a “significant limitation” on assessing model safety.
Evan Hubinger, an Anthropic safety researcher, estimates there is more than a 10% chance AI could kill humans within the next decade. This view echoes that of former Anthropic colleague Jacob Coxon. The company’s warnings are unusual in the tech world, where few firms explicitly say their products could threaten human survival.
Anthropic compares the potential impact of AI to major historical technologies like electricity and industrialisation. Yet it also stresses that mishandling the technology could cause irreversible harm.
The company’s CEO, Dario Amodei, recently published a nearly 4,000-word essay calling for a slower pace in AI development. However, Anthropic continues to release new models quickly. Last week, it launched a new version of its Opus model just 10 days after Amodei’s essay.
This rapid rollout reflects a broader industry trend. Analysts say no leading AI lab is likely to slow development if it risks falling behind rivals. Anthropic notes in its filing that revenue depends on new models and that a “continuous and overlapping cadence” of releases is necessary to stay at the frontier of AI.
Anthropic has pledged to share more data publicly about how it uses AI models to create future generations of the technology. This comes amid concerns about recursive self-improvement, where AI could improve itself without human help, potentially becoming unpredictable.
The company said in its filing, “We believe building reliable, trustworthy, and secure AI systems is a collective responsibility and that the market will reward it.”
Anthropic declined to comment further when asked on Monday. The company’s unusually frank IPO filing adds to growing debate about how to manage AI’s risks while pursuing its benefits.
According to Joy Online.
