arXiv:2503.11074cs.AIcs.CL2025-03被引 7

对比大模型与推理模型在智能体中的表现,发现各有所长。

Exploring the Necessity of Reasoning in LLM-based Agent Scenarios

  • 用九项任务测试大模型与推理模型的协作能力
  • 推理模型擅长规划类任务,执行模型更高效
  • 混合配置能兼顾速度与深度,适合复杂决策

大型推理模型(LRMs)的兴起标志着计算推理能力的重大进步。然而,这一进展打破了传统以执行为导向的大型语言模型(LLMs)为根基的智能体框架。为此,我们提出了LaRMA框架,涵盖工具使用、规划设计和问题解决三类共九项任务,评估了三种顶级LLMs(如Claude3.5-sonnet)和五种领先LRMs(如DeepSeek-R1)。研究回答四个核心问题:在规划设计等高推理需求任务中,LRMs凭借迭代反思优于LLMs;在工具使用等执行导向任务中,LLMs更具效率;将LLM作为执行者、LRM作为反思者组成的混合架构可实现性能最优;但LRMs增强推理能力也带来更高计算开销、更长处理时间及过思考、忽视事实等行为问题。本研究推动对深度思考与过度思考之间平衡的深入探讨,为未来智能体设计奠定关键基础。

原文摘要 · Abstract (English)

The rise of Large Reasoning Models (LRMs) signifies a paradigm shift toward advanced computational reasoning. Yet, this progress disrupts traditional agent frameworks, traditionally anchored by execution-oriented Large Language Models (LLMs). To explore this transformation, we propose the LaRMA framework, encompassing nine tasks across Tool Usage, Plan Design, and Problem Solving, assessed with three top LLMs (e.g., Claude3.5-sonnet) and five leading LRMs (e.g., DeepSeek-R1). Our findings address four research questions: LRMs surpass LLMs in reasoning-intensive tasks like Plan Design, leveraging iterative reflection for superior outcomes; LLMs excel in execution-driven tasks such as Tool Usage, prioritizing efficiency; hybrid LLM-LRM configurations, pairing LLMs as actors with LRMs as reflectors, optimize agent performance by blending execution speed with reasoning depth; and LRMs' enhanced reasoning incurs higher computational costs, prolonged processing, and behavioral challenges, including overthinking and fact-ignoring tendencies. This study fosters deeper inquiry into LRMs' balance of deep thinking and overthinking, laying a critical foundation for future agent design advancements.

大模型智能体推理能力混合架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。