arXiv:2509.23537cs.AI2025-09被引 9

多智能体协作比单个大模型更准,能自动达成共识。

Beyond the Strongest LLM: Multi-Turn Multi-Agent Orchestration vs. Single LLMs on Benchmarks

  • 多个大模型通过多轮投票协作,逐步达成一致答案。
  • 在三大测试集上,协作效果超越最强单模型。
  • 观察他人投票会加速共识但易导致过早定论。

我们研究了多轮多智能体协作,即多个大语言模型(LLM)通过多轮迭代提出答案或投票,直至达成共识。在GPQA-Diamond、IFEval和MuSR三个数据集上,使用Gemini 2.5 Pro、GPT-5、Grok 4和Claude Sonnet 4四种模型进行实验:(i) 对比协作与单模型基线;(ii) 在GPQA-Diamond上分析作者身份可见性及投票过程可见性的影响。结果表明,协作方案达到或超过最强单模型表现,并持续优于其他模型。最佳协作性能分析显示仍有提升空间。消融实验发现,公开作者身份会增加自投和并列票数,而展示实时投票会加剧从众效应,虽加快收敛,但可能引发过早共识。

原文摘要 · Abstract (English)

We study multi-turn multi-agent orchestration, where multiple large language model (LLM) agents interact over multiple turns by iteratively proposing answers or casting votes until reaching consensus. Using four LLMs (Gemini 2.5 Pro, GPT-5, Grok 4, and Claude Sonnet 4) on GPQA-Diamond, IFEval, and MuSR, we conduct two experiments: (i) benchmarking orchestration against single-LLM baselines; and (ii) ablations on GPQA-Diamond that vary whether agents see who authored answers and whether they can observe ongoing votes. Orchestration matches or exceeds the strongest single model and consistently outperforms the others. Analysis of best-achievable orchestration performance shows potential for further gains. The ablations show that revealing authorship increases self-voting and ties, and that showing ongoing votes amplifies herding, which speeds convergence but can sometimes yield premature consensus.

多智能体协同推理大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。