让AI辩论更有效:用观点多样性和可信度沟通提升决策质量
Demystifying Multi-Agent Debate: The Role of Confidence and Diversity
- 引入观点多样性初始化,提高正确答案初始出现概率
- 通过可信度调节辩论过程,使结果稳定趋向正确答案
- 在6个推理问答数据集上超越传统辩论和多数投票
多智能体辩论(MAD)通过测试时扩展提升大语言模型性能,但研究表明,原始MAD常因计算成本高而表现不如简单多数投票。已有研究指出,在同质代理与均匀信念更新下,辩论无法可靠改进结果。借鉴人类协商与集体决策的研究,本文识别出两个关键缺失机制:(i) 初始观点多样性,(ii) 显式且校准的可信度表达。为此提出两种轻量级干预:首先,多样性感知初始化,从更丰富的候选答案中选择初始立场,提高正确假设的初始覆盖率;其次,可信度调节辩论协议,允许代理以校准方式表达信心,并根据他人可信度调整自身信念。理论上证明,多样性初始化可提升成功概率而不改变更新动态;可信度调节则使辩论系统性地收敛至正确结论。实证上,在六个推理型问答基准上,该方法持续优于原始MAD与多数投票。研究将人类协商机制与基于LLM的辩论联系起来,表明简单、原则性的改进可显著增强辩论有效性。
原文摘要 · Abstract (English)
Multi-agent debate (MAD) is widely used to improve large language model (LLM) performance through test-time scaling, yet recent work shows that vanilla MAD often underperforms simple majority vote despite higher computational cost. Studies show that, under homogeneous agents and uniform belief updates, debate preserves expected correctness and therefore cannot reliably improve outcomes. Drawing on findings from human deliberation and collective decision-making, we identify two key mechanisms missing from vanilla MAD: (i) diversity of initial viewpoints and (ii) explicit, calibrated confidence communication. We propose two lightweight interventions. First, a diversity-aware initialisation that selects a more diverse pool of candidate answers, increasing the likelihood that a correct hypothesis is present at the start of debate. Second, a confidence-modulated debate protocol in which agents express calibrated confidence and condition their updates on others' confidence. We show theoretically that diversity-aware initialisation improves the prior probability of MAD success without changing the underlying update dynamics, while confidence-modulated updates enable debate to systematically drift to the correct hypothesis. Empirically, across six reasoning-oriented QA benchmarks, our methods consistently outperform vanilla MAD and majority vote. Our results connect human deliberation with LLM-based debate and demonstrate that simple, principled modifications can substantially enhance debate effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。