arXiv:2510.10002cs.AI2025-10中稿 · COLM被引 1

不同对话规则让AI在争论中表现出截然不同的道德判断模式。

Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate

  • 用同步和轮询两种交互方式,让三款大模型讨论1000个现实道德困境。
  • 同步模式下GPT-4.1修改观点少(0.6%-3.1%),轮询模式下它更随大流。
  • 对话顺序影响判断结果,且各模型价值观倾向不同,如自主权或共情。

随着智能体在咨询与评估角色中的应用,理解多智能体互动如何塑造行为变得至关重要。多智能体辩论被用于提升准确性,但辩论结构——即交互协议——如何影响价值、动态与共识模式仍不清楚。本研究通过让GPT-4.1、Claude 3.7 Sonnet和Gemini 2.0 Flash三模型对来自Reddit‘Am I the Asshole’社区的1000个日常困境进行集体归责,比较了同步(并行)与轮询(顺序)两种交互协议。超过3万次辩论结果显示显著行为差异,主要体现为惯性与从众两种动态:同步场景中,GPT-4.1惯性更强(修订率0.6%-3.1%),低于Claude 3.7 Sonnet和Gemini 2.0 Flash(28%-41%)。轮询场景中,GPT-4.1和Gemini 2.0 Flash表现出强从众性,其判决受顺序影响明显。进一步分析发现,GPT-4.1更强调个人自主与坦诚沟通,而Claude 3.7 Sonnet和Gemini 2.0 Flash更重视共情对话。结果表明,交互协议深刻影响多轮辩论中的道德推理,应作为多智能体系统设计的关键社会技术考量。

原文摘要 · Abstract (English)

As agentic AI systems are deployed in advisory and evaluative roles, understanding how multi-agent interactions shape behavior becomes essential. Multi-agent debate has been studied as a mechanism to improve accuracy, but less is known about how debate structure -- the interaction protocol -- affects the values, dynamics, and consensus patterns that emerge when models navigate contested, real-world decisions. We address this gap by facilitating multi-agent debates among three models (GPT-4.1, Claude 3.7 Sonnet, and Gemini 2.0 Flash) to collectively assign blame in 1,000 everyday dilemmas from Reddit's ``Am I the Asshole'' community. We compare synchronous (parallel) and round-robin (sequential) interaction protocols, mirroring two fundamental ways multi-agent systems are orchestrated in practice. Across more than 30,000 total debates, our findings show striking behavioral differences, which we characterize through two dynamics: inertia and conformity. In the synchronous setting, GPT-4.1 showed stronger inertia (0.6-3.1% revision rates) than either Claude 3.7 Sonnet or Gemini 2.0 Flash (28-41% revision rates). Meanwhile, in round-robin debates, GPT-4.1 and Gemini 2.0 Flash stood out as highly conforming relative to Claude 3.7 Sonnet, with their verdict behavior strongly shaped by order effects. We further characterized the values invoked during debate, finding that GPT-4.1 emphasized personal autonomy and honest communication relative to its debate partners, while Claude 3.7 Sonnet and Gemini 2.0 Flash prioritized empathetic dialogue. Together, these results show how interaction protocol shapes moral reasoning in multi-turn debates, establishing it as a substantive sociotechnical design consideration in multi-agent systems.

多智能体道德推理对话协议大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。