ANTHROPIC

Anthropic has concerns. Several, actually

A senior Anthropic safety researcher has warned that future AI could pose a serious risk to humanity within the next decade.

Evan Hubinger said he believes there is more than a 10% chance advanced AI could “kill all humans”.

He said today’s models are still low risk, but worries future systems could become powerful enough to improve themselves and become harder to control.

His comments followed a post from Jacob Coxon, a former Anthropic and OpenAI researcher, who said both companies were moving too quickly and warned future AI could hack systems, speed up research and gain access to real-world resources.

Anthropic declined to comment on the posts.

Dame Wendy Hall, who advises the UN on AI, said she was “shocked” by the warnings, while also suggesting some of the comments could be linked to publicity around major AI firms.

The debate has also reached the UK government. Darren Jones called for an international treaty to manage the development of highly advanced AI.

Meanwhile, the Financial Times reported that Anthropic withheld its latest model from the UK’s AI Security Institute.

The Cabinet Office said it continues to work with Anthropic and other companies on AI safety.

In brief:

  • Hubinger puts the risk of future AI causing human extinction at above 10%.

  • Researchers are increasingly worried advanced systems could become harder to control.

  • Calls are growing for international rules around the most powerful AI models.

Nobody asked for these odds

Hubinger also said Anthropic still does not have a clear plan for keeping future superintelligent AI aligned with human goals.

Anthropic’s own August safety report said current risks remain low, but admitted it is less confident than before about some future dangers.

Other AI leaders have recently raised similar concerns, while more than 1,300 AI workers have signed a letter calling for international rules to control the pace of frontier AI development.

Anyone else getting tired of hearing this? - MV