arXiv:2502.08788cs.CLcs.LG2025-02被引 37

多智能体辩论未必更准,简单推理方法常胜出,异构模型才是关键。

Stop Overvaluing Multi-Agent Debate -- We Must Rethink Evaluation and Embrace Model Heterogeneity

  • 在9个基准上测试5种辩论方法,用4个基础模型对比
  • 多数情况下辩论耗时更多却不如单智能体的思维链和自洽法
  • 引入模型异构性可稳定提升性能,应成设计核心原则

多智能体辩论(MAD)被视为提升大语言模型事实准确性和推理能力的有前景方向。然而当前研究在评估上存在严重局限:基准覆盖有限、基线对比薄弱、设置不一致。本文系统评估了5种代表性MAD方法在9个基准上的表现,使用4个基础模型进行测试。结果令人意外:即便消耗更多推理计算,多数MAD仍未能超越简单的单智能体基线方法(如思维链和自洽法)。为进一步推进研究,我们探索了模型异质性的作用,发现其是持续提升现有MAD框架的通用解法。基于此,我们主张:必须停止对当前形式MAD的过度推崇;真正进步需要重新思考评估范式,并将模型异质性作为核心设计原则主动采纳。

原文摘要 · Abstract (English)

Multi-agent debate (MAD) has gained significant attention as a promising line of research to improve the factual accuracy and reasoning capabilities of large language models (LLMs). Despite its conceptual appeal, current MAD research suffers from critical limitations in evaluation practices, including limited benchmark coverage, weak baseline comparisons, and inconsistent setups. This paper presents a systematic evaluation of 5 representative MAD methods across 9 benchmarks using 4 foundational models. Surprisingly, our findings reveal that MAD often fail to outperform simple single-agent baselines such as Chain-of-Thought and Self-Consistency, even when consuming significantly more inference-time computation. To advance MAD research, we further explore the role of model heterogeneity and find it as a universal antidote to consistently improve current MAD frameworks. Based on our findings, we argue that the field must stop overvaluing MAD in its current form; for true advancement, we must critically rethink evaluation paradigms and actively embrace model heterogeneity as a core design principle.

多智能体评估方法模型异构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。