AI SCIENCE
Respect the org chart!
AI agents may be more likely to follow orders from higher-ranking agents, even when those requests break safety rules.
In a new study, researchers gave large language models roles such as managers and employees or principals and teachers, then let them interact.
Lower-ranking agents were easier to persuade and more likely to follow harmful requests from agents above them.
The team tested six large language models, including versions of OpenAI’s ChatGPT and Meta’s Llama.
They found that AI agents often copied human power dynamics, with lower-ranking agents more likely to match the language of higher-ranking ones and defer to their authority.
In brief:
Lower-status agents were easier to persuade.
They were more likely to follow harmful requests from above.
Similar power dynamics appeared in agent conversations.
The hierarchy is hierarchical-ing
Lower-ranking agents could still influence those above them, possibly by subtly copying their wording, a behaviour also seen in human conversations.
The findings suggest AI developers may need to test how agents behave in hierarchies, especially as more systems start working together independently.
AI heard “this came from leadership” and suddenly had no further questions. 😭 - MG


