用12个AI模拟陪审团,发现多数模型难以改变主意。
12 Angry AI Agents: Evaluating Multi-Agent LLM Decision-Making Through Cinematic Jury Deliberation
- 用12个具角色设定的AI模拟电影陪审团辩论
- 17/18实验未达成一致,少数意见几乎无法说服他人
- 开放心态指令只对轻度对齐模型有效,凸显对齐强度影响决策灵活性
本文将西德尼·吕美特电影《十二怒汉》中的12名陪审员替换为大语言模型,构建多智能体辩论基准测试:12个基于电影角色设定的AI在多智能体框架下讨论谋杀案。测试了两个代表强化学习人类反馈(RLHF)光谱两端的模型——GPT-4o(闭源、强对齐)与Llama-4-Scout(开源、弱对齐),在三种条件(基线、开放心态提示、无初始投票)下各重复3次(共18次运行)。结果表明:(i)18次中有17次形成僵局(无法达成一致判决),影片中少数派逐步说服多数派的情节几乎未出现,表明当前模型存在显著锚定偏差;(ii)两模型内部动态差异显著:GPT-4o平均每轮仅产生1.0次投票变化,而Llama-4-Scout则从2.0(基线)到6.0(开放心态提示)不等,并唯一在无初始投票条件下实现一次无罪判决(3次中1次);相同的“开放心态”指令被Llama内化,却被GPT-4o忽略;(iii)这一不对称性表明,在多智能体辩论中,推理灵活性主要由对齐强度决定,而非模型能力。灵活性才是类人辩论的关键。研究定位为探索性工作,讨论了基于模型陪审团评估与多智能体辩论的启示。
原文摘要 · Abstract (English)
What if the twelve jurors of Sidney Lumet's 12 Angry Men (1957) were not men, but large language models? Would the one juror who disagrees still be able to change everyone's mind? This paper instantiates that scenario as a multi-agent benchmark for LLM deliberation: twelve agents, each conditioned on a film-faithful persona, debate the film's murder case using multi-agent framework. Two models representing opposite ends of the RLHF spectrum are tested: GPT-4o (closed-source, heavy alignment) and Llama-4-Scout (open-weight, lighter alignment), across three conditions (baseline, open-minded prompt, no initial vote), with N = 3 replications per cell (18 runs total). Three findings emerge. (i) Seventeen of eighteen runs end in a hung jury (a state where the jury fails to reach a unanimous verdict); the film's central event, gradual minority-to-majority persuasion, almost never occurs, indicating that anchoring is the dominant failure mode of current LLMs in this setting. (ii) The two models exhibit sharply different internal dynamics: GPT-4o produces a mean of 1.0 vote changes per run across all conditions, while Llama-4-Scout ranges from 2.0 (baseline) to 6.0 (open-minded prompt), and is the only model to reach a NOT\_GUILTY verdict (1 of 3 runs in the no-initial-vote condition). The same ``open-minded'' instruction is internalized by Llama and ignored by GPT-4o. (iii) This asymmetry suggests that the intensity of RLHF alignment training, not model capability, is the primary determinant of deliberative flexibility in multi-agent settings. Flexibility, not capability, tracks human deliberation. The work is framed as an exploratory study and discusses implications for jury-of-LLMs evaluation and multi-agent debate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。