多智能体协作让大模型更倾向功利主义决策
Many LLMs Are More Utilitarian Than One
- 让多个大模型讨论道德困境,发现群体决策更易接受伤害少数人以造福多数
- 所有模型在群体中对道德违规的接受度提升,最大增幅达27%
- 适合关注AI伦理、多智能体系统设计的研究者阅读
道德判断是大语言模型社会推理的核心。随着多智能体系统日益重要,理解模型在协作中与独立运作时的表现差异变得关键。人类在群体讨论中会表现出功利主义增强:更倾向于接受造成伤害但最大化整体利益的行为。我们测试了六种模型在两种条件下的表现:(1) 独立推理(Solo),(2) 成对或三人组多轮讨论(Group)。在个人困境中,所有模型在群体条件下均更认可道德违规行为,表现出与人类相似的功利主义增强效应。但机制不同:人类因更关注结果而变得更功利,而大模型群体则表现为对规范敏感性下降或更强的公正性。我们报告了不同模型在何时及多强程度上出现该效应,并讨论了提示词与智能体组合对效果的影响。最后探讨了对人工智能对齐、多智能体设计及人工道德推理的启示。代码已开源。
原文摘要 · Abstract (English)
Moral judgment is integral to large language models' (LLMs) social reasoning. As multi-agent systems gain prominence, it becomes crucial to understand how LLMs function when collaborating compared to operating as individual agents. In human moral judgment, group deliberation leads to a Utilitarian Boost: a tendency to endorse norm violations that inflict harm but maximize benefits for the greatest number of people. We study whether a similar dynamic emerges in multi-agent LLM systems. We test six models on well-established sets of moral dilemmas across two conditions: (1) Solo, where models reason independently, and (2) Group, where they engage in multi-turn discussions in pairs or triads. In personal dilemmas, where agents decide whether to directly harm an individual for the benefit of others, all models rated moral violations as more acceptable when part of a group, demonstrating a Utilitarian Boost similar to that observed in humans. However, the mechanism for the Boost in LLMs differed: While humans in groups become more utilitarian due to heightened sensitivity to decision outcomes, LLM groups showed either reduced sensitivity to norms or enhanced impartiality. We report model differences in when and how strongly the Boost manifests. We also discuss prompt and agent compositions that enhance or mitigate the effect. We end with a discussion of the implications for AI alignment, multi-agent design, and artificial moral reasoning. Code available at: https://github.com/baltaci-r/MoralAgents
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。