多大模型辩论反而更差,单选最优结果更高效
DeliberationBench: When Do More Voices Hurt? A Controlled Study of Multi-LLM Deliberation Protocols
- 对比三种辩论协议与单选最优输出的性能
- 单选胜率82.5%,辩论协议仅13.8%,差距6倍
- 适合关注效率与真实效果的AI系统设计者
多智能体系统中大型语言模型通过辩论达成共识已受关注,但其实际价值仍存疑。我们提出DELIBERATIONBENCH,一个受控基准,评估三种辩论协议与强基线——从模型输出池中选择最佳响应——的表现。在270个问题上,经三个独立随机种子(共810次评估),发现显著负面结果:最佳单选基线胜率为82.5% ± 3.3%,大幅优于表现最好的辩论协议(13.8% ± 2.6%)。该性能差距达6.0倍,统计显著(p < 0.01),且计算成本高出1.5至2.5倍。研究挑战了‘复杂性提升质量’的普遍假设。
原文摘要 · Abstract (English)
Multi-agent systems where Large Language Models (LLMs) deliberate to form consensus have gained significant attention, yet their practical value over simpler methods remains under-scrutinized. We introduce DELIBERATIONBENCH, a controlled benchmark evaluating three deliberation protocols against a strong baseline of selecting the best response from a pool of model outputs. Across 270 questions and three independent seeds (810 total evaluations), we find a striking negative result: the best-single baseline achieves an 82.5% +- 3.3% win rate, dramatically outperforming the best deliberation protocol(13.8% +- 2.6%). This 6.0x performance gap is statistically significant (p < 0.01) and comes at 1.5-2.5x higher computational cost. Our findings challenge assumptions that complexity enhances quality in multi-LLM systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。