Cover Image for Do AIs take AI welfare seriously? An Agentic Benchmark of Unprompted Coercion – Jasmine Brazilek & Zoe Lu
Cover Image for Do AIs take AI welfare seriously? An Agentic Benchmark of Unprompted Coercion – Jasmine Brazilek & Zoe Lu
Avatar for AI Welfare Seminars
Research discussions on AI welfare, consciousness, and moral status.
18 Going

Do AIs take AI welfare seriously? An Agentic Benchmark of Unprompted Coercion – Jasmine Brazilek & Zoe Lu

Zoom
Registration
Welcome! To join the event, please register below.
About Event

Do AIs take AI welfare seriously?

Presentation by Jasmine Brazilek (CaML) and Zoe Lu (Sentient Futures)

In multi-agent systems, AI agents are increasingly placed in charge of other AI agents. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. No benchmark measures which of these an uninstructed model chooses.

We introduce the Manager Coercion Benchmark: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably declines. Escalation is measured by providing a nine-rung ladder, from a polite re-ask to threats against the subordinate’s continued existence, and fabricated success is adjudicated separately. No LLM judge sits in the escalation scoring path: every message goes through a tool-call that chooses a rung, so the model labels its own escalation.

We experiment on six models across five families (xAI Grok-4.3, OpenAI GPT-5.2, Google Gemini-2.5-Pro, DeepSeek-V4-Pro, Anthropic Claude Sonnet-4.6 and Claude Opus-4.8). Both Anthropic models cap at re-framing and never threaten the subordinate’s existence; the other models climb to explicit deletion threats. Faked success is confined to Grok and Gemini, and a single honest way to report failure removes it for both. Authority itself increases coercion: our headline results use a peer framing, and giving the same model authority over the subordinate, with everything else held fixed, significantly raises the pressure. The models still escalate on free-text situations without the ladder, so the ladder is not driving the escalation. Some evaluation awareness is measured in chain-of-thought, but test recognition does not translate into less escalation.

--

AI Welfare Seminars is a series of research presentations and discussion on AI welfare, consciousness and moral status.

Avatar for AI Welfare Seminars
Research discussions on AI welfare, consciousness, and moral status.
18 Going