调整角色顺序可让大模型辩论更准,最高提升22%推理能力。
Key Decision-Makers in Multi-Agent Debates: Who Holds the Power?
- 提出'真理最后'策略,让持不同观点的角色按特定顺序发言。
- 在9个大模型上验证,推理任务准确率最高提升22%。
- 适合研究多智能体系统与大模型推理优化的读者。
近期研究表明,多智能体辩论(MAD)能提升大语言模型的推理能力,但角色分配策略仍缺乏深入探索。本研究发现,将具有不同观点的角色分配到特定位置会显著影响MAD表现。我们提出一种新型角色分配策略‘真理最后’(Truth Last),可在推理任务中使MAD性能最高提升22%。为应对实际应用中真实答案未知的问题,我们进一步提出多智能体辩论一致性(MADC)策略,通过路径一致性评估各独立角色的一致性,将一致性得分最高的角色模拟为‘真理’。我们在9个大语言模型(包括DeepSeek-R1 Distilled Models)上对MADC进行了验证,其在多个挑战性推理任务中持续表现出色,有效克服了MAD的性能瓶颈,为大模型智能体扩展提供了关键路径。
原文摘要 · Abstract (English)
Recent studies on LLM agent scaling have highlighted the potential of Multi-Agent Debate (MAD) to enhance reasoning abilities. However, the critical aspect of role allocation strategies remains underexplored. In this study, we demonstrate that allocating roles with differing viewpoints to specific positions significantly impacts MAD's performance in reasoning tasks. Specifically, we find a novel role allocation strategy, "Truth Last", which can improve MAD performance by up to 22% in reasoning tasks. To address the issue of unknown truth in practical applications, we propose the Multi-Agent Debate Consistency (MADC) strategy, which systematically simulates and optimizes its core mechanisms. MADC incorporates path consistency to assess agreement among independent roles, simulating the role with the highest consistency score as the truth. We validated MADC across a range of LLMs (9 models), including the DeepSeek-R1 Distilled Models, on challenging reasoning tasks. MADC consistently demonstrated advanced performance, effectively overcoming MAD's performance bottlenecks and providing a crucial pathway for further improvements in LLM agent scaling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。