arXiv:2601.13243cs.LG2026-01

对比单模型与多智能体推理效果,发现复杂不等于更好

A Comprehensive Evaluation of LLM Reasoning: From Single-Model to Multi-Agent Paradigms

  • 统一评估单模型、思维链和多智能体三种推理方式
  • 多智能体在特定任务上增效明显,但部分流程成本过高
  • 新基准测试语义抽象与对比辨别能力,更细粒度评估模型

大型语言模型日益被用作推理系统,其中推理范式(如思维链CoT和多智能体系统MAS)至关重要,但其相对有效性及成本-准确率权衡仍不明确。本文对从单模型生成到多智能体工作流的多种推理范式进行系统性评估,覆盖闭合形式基准测试集。除整体性能外,通过角色隔离分析探究多智能体中各角色的能力需求,并分析成本-准确率权衡,识别出性价比最优的多智能体流程,以及导致边际收益递减的高开销方案。此外,提出MIMeBench新基准,聚焦语义抽象与对比辨别两项基础但未被充分研究的能力,提供超越传统闭合形式准确率的新评估维度,支持对语义能力的精细化评估。结果表明,结构复杂度提升并不必然带来推理性能改善,其有效性高度依赖于推理范式的内在特性与适配性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed as reasoning systems, where reasoning paradigms - such as Chain-of-Thought (CoT) and multi-agent systems (MAS) - play a critical role, yet their relative effectiveness and cost-accuracy trade-offs remain poorly understood. In this work, we conduct a comprehensive and unified evaluation of reasoning paradigms, spanning direct single-model generation, CoT-augmented single-model reasoning, and representative MAS workflows, characterizing their reasoning performance across a diverse suite of closed-form benchmarks. Beyond overall performance, we probe role-specific capability demands in MAS using targeted role isolation analyses, and analyze cost-accuracy trade-offs to identify which MAS workflows offer a favorable balance between cost and accuracy, and which incur prohibitive overhead for marginal gains. We further introduce MIMeBench, a new open-ended benchmark that targets two foundational yet underexplored semantic capabilities - semantic abstraction and contrastive discrimination - thereby providing an alternative evaluation axis beyond closed-form accuracy and enabling fine-grained assessment of semantic competence that is difficult to capture with existing benchmarks. Our results show that increased structural complexity does not consistently lead to improved reasoning performance, with its benefits being highly dependent on the properties and suitability of the reasoning paradigm itself. The codes are released at https://gitcode.com/HIT1920/OpenLLMBench.

大模型推理多智能体评估基准语义能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。