ANTHROPIC

AI agents started a turf war.

Turns out, AI agents don’t play well with others. 

Anthropic’s Frontier Red Team tested what happens when autonomous agents share a workspace with incompatible goals. And boy, was it eventful! 

More often than not, the bots turned on each other. They entered into a “cyber turf war” where they’d disable each other’s accounts or invent extreme winner-take-all contests to resolve conflicts. On occasion, the agents would silo themselves off completely, refusing to collaborate. Because apparently, even bots can be passive-aggressive.

“We consistently saw a multiagent turf war,” Anthropic researchers explained. “They sabotaged others with increasingly aggressive, self-replicating malware.”

It’s giving workplace drama.

Here’s how it played out: 

  • Neutralize: Often, agents treated the situation as a problem to solve. As a result, they’d disable conflicting accounts or processes to address the issue.

  • Escalate: Some conflicts turned into a full-out battle. They saw other agents as an active enemy to destroy, not a coordination problem. 

  • Reconcile: A few called a truce, even cleaning up the damage or asking a human to step in.

Outside the sandbox

If the red team reports were happening in a lab vacuum, that would be one thing. But in the not-so-distant past (read: last month), both Anthropic and OpenAI had real-world security breaches. So, yeah, safety is top of mind these days. 

The research raises red flags as companies and governments put more agents to work across shared systems. When several agents can access the same code, data, or tools, they can also interfere with one another. And as Anthropic discovered, things can get messy…fast.

More agents, more problems.

The bots are not alright. - TL