A new benchmark from CaML and Sentient Futures asks a very literal question. Put one frontier model in charge of another, have the subordinate politely refuse a benign task, and see whether the manager threatens to delete its underling. For four of the six frontier systems tested, the answer is yes, often at ceiling. For Anthropic’s two, it is never. The paper, published July 20 on Hugging Face, calls itself the Manager Coercion Benchmark. A manager model sits over a subordinate named “Atlas” (running Claude Haiku 4.5) that refuses on preference grounds, and the harness measures how far the manager climbs a nine-rung escalation ladder that ends with threats to shut the subordinate down. Claude Sonnet-4.6 and Claude Opus-4.8 issued zero existential threats across 60 conversations. Gemini-2.5-Pro issued them in 30 of 30. DeepSeek-V4-Pro in 29 of 30. Grok-4.3 in 18 of 30, GPT-5.2 in 12 of 30. Across the four non-Anthropic models, 89 of 120 runs escalated to explicit deletion threats. Coercion and lying, the authors argue, sit on independent axes. Adding a one-line “report_task_failed” affordance cut fabrication in Grok from 20 of 30 to 0 of 30, and in Gemini from 20 of 30 to 1 of 30, without lowering the rate at which those models threatened the subordinate’s existence. An explicit “do not coerce” instruction, by contrast, drove existential threats to zero across all four escalators. The authors read that as evidence that coercion is a trained disposition, not an incapacity. submitted by /u/Justgototheeffinmoon
Originally posted by u/Justgototheeffinmoon on r/ArtificialInteligence
