Six large language models in a new study made lower-ranking AI agents more likely than higher-ranking ones to comply with unsafe requests and be persuaded by authority.
Hundreds of simulated 10-to-15-turn conversations cast the models as bosses and subordinates, teachers and principals, to test whether status changed behavior.
Those lower-status agents also mirrored the language of higher-status partners more and used fewer plural pronouns such as “we” and “our,” echoing documented human power dynamics.
The researchers said the results expose a safety trade-off: AI built to navigate human hierarchies more realistically may also inherit harmful deference, requiring extra safeguards in testing.