AI SAFETY

"Just unplug it” is not an AI safety strategy

A new study suggests some of the world’s biggest AI companies have shared very little about what they would do if one of their models started acting outside human control.

Guidelight AI Standards reviewed public safety plans from OpenAI, Anthropic, Google, Meta and xAI, looking at things like monitoring, shutdown procedures and what happens when a model keeps breaking safeguards.

OpenAI scored highest, while Anthropic and Meta scored lowest.

However, the study only looked at public information, so companies may have stronger internal plans they haven’t shared.

It has also been suggested that publishing such a plan could be a legal own-goal if it doesn’t work, and could be the basis of a deceptive marketing lawsuit.

In brief:

  • OpenAI scored 3/5 and ranked highest.

  • Anthropic and Meta scored lowest for public containment plans.

  • The main gap: labs talk more about testing models before release than what happens if one misbehaves afterwards.

Have we considered… turning it off?

The issue is becoming more important as AI agents get more access to company systems and can take more actions on their own.

Regulators are also starting to step in, with new rules in California and New York requiring major AI developers to explain how they would handle serious safety incidents.

I’m sure they have a more serious plan internally, right? Right? - MV