让不同模型辩论可激发更强推理能力,超越单一大模型。
Diversity of Thought Elicits Stronger Reasoning Capabilities in Multi-Agent Debate Frameworks
- 用多个不同模型进行辩论,比相同模型更有效。
- 4轮辩论后,中等模型组合在GSM-8K上达91%准确率,超GPT-4。
- 适合追求高推理性能、探索多智能体协作的研究者。
大型语言模型在自然语言生成方面表现优异,但在数学推理等任务中常自信地输出错误答案。链式思维提示、自我验证及多智能体辩论是提升模型推理与事实准确性的策略。基于Du等人提出的多智能体辩论框架,我们发现该方法在任何模型规模下均有效,且思维多样性能显著增强推理能力。在不同模型规模下,使用多样化的训练模型进行辩论时,数学推理表现最优。令人瞩目的是,在4轮辩论后,由Gemini-Pro、Mixtral 7BX8和PaLM 2-M组成的多样化中等容量模型集合在GSM-8K基准上达到91%准确率,优于GPT-4;而使用3个Gemini-Pro实例仅达82%。此外,该组合在ASDiv基准上实现94%新纪录。结果表明,未来AI的发展方向是代理化,多样化协作智能体可涌现出超越最强单一模型的能力。
原文摘要 · Abstract (English)
Large language models (LLMs) excel in natural language generation but often confidently produce incorrect responses, especially in tasks like mathematical reasoning. Chain-of-thought prompting, self-verification, and multi-agent debate are among the strategies proposed to improve the reasoning and factual accuracy of LLMs. Building on Du et al.'s multi-agent debate framework, we find that multi-agent debate helps at any model scale, and that diversity of thought elicits stronger reasoning in debating LLMs. Across various model sizes, performance on mathematical reasoning tasks benefits most when diverse trained models are used. Remarkably, after 4 rounds of debate, a diverse set of medium-capacity models (Gemini-Pro, Mixtral 7BX8, and PaLM 2-M) outperforms GPT-4 on the GSM-8K benchmark, scoring 91% accuracy. By comparison, when 3 instances of Gemini-Pro are used, performance only reaches 82%. Finally, this diverse set of medium-capacity models sets a new state-of-the-art performance on the ASDiv benchmark (94%). These results underscore the idea that the future of AI is agentic, with diverse cooperating agents yielding emergent capabilities beyond even the most powerful individual models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。