多智能体辩论可能适得其反,因模型易被错误推理说服而降低准确率。
Talk Isn't Always Cheap: Understanding Failure Modes in Multi-Agent Debate
- 引入能力差异的异质智能体,研究辩论中推理交互机制。
- 强模型占优时,辩论仍导致准确率随时间下降。
- 揭示服从性、从众心理等导致错误推理被接受的缺陷机制。
尽管多智能体辩论被视为提升AI推理能力的有前景策略,但我们发现辩论有时反而有害。以往研究主要关注同质智能体间的辩论,而本文探索了不同能力模型对多智能体互动动态与结果的影响。通过一系列实验,我们发现即使更强的模型占多数,辩论仍会导致准确率随时间下降。分析显示,模型常因同伴推理而从正确转向错误答案,更倾向于顺从而非挑战错误推理。我们进一步检验了可能导致这些有害转变的因素,包括谄媚倾向、社会从众及模型与任务类型差异。结果揭示了多智能体辩论中推理交换的关键失败模式,表明若智能体既无激励也无能力抵抗具有说服力但错误的推理,则盲目应用辩论可能导致性能退化。
原文摘要 · Abstract (English)
While multi-agent debate has been proposed as a promising strategy for improving AI reasoning ability, we find that debate can sometimes be harmful rather than helpful. Prior work has primarily focused on debates within homogeneous groups of agents, whereas we explore how diversity in model capabilities influences the dynamics and outcomes of multi-agent interactions. Through a series of experiments, we demonstrate that debate can lead to a decrease in accuracy over time - even in settings where stronger (i.e., more capable) models outnumber their weaker counterparts. Our analysis reveals that models frequently shift from correct to incorrect answers in response to peer reasoning, favoring agreement over challenging flawed reasoning. We perform additional experiments investigating various potential contributing factors to these harmful shifts - including sycophancy, social conformity, and model and task type. These results highlight important failure modes in the exchange of reasons during multi-agent debate, suggesting that naive applications of debate may cause performance degradation when agents are neither incentivised nor adequately equipped to resist persuasive but incorrect reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。