arXiv:2601.12538cs.AIcs.CL2026-01被引 38

让大模型像人一样自主规划、行动并学习,突破传统推理局限。

Agentic Reasoning for Large Language Models

  • 将大模型视为能自主规划与学习的智能体,分三层构建能力。
  • 通过反馈与记忆实现自我进化,支持多智能体协作与共享目标。
  • 适合研究智能体系统、自主决策及复杂任务应用的开发者。

推理是推断、问题解决和决策的基础认知过程。尽管大语言模型(LLMs)在封闭环境中有强大推理能力,但在开放且动态的环境中表现不佳。代理式推理标志着范式转变:将LLMs重新定义为能通过持续交互进行规划、行动和学习的自主智能体。本文从三个互补维度组织代理式推理:基础层,建立单智能体在稳定环境中的核心能力,包括规划、工具使用和搜索;自进化层,研究智能体如何通过反馈、记忆和适应性优化这些能力;协同多智能体层,将智能扩展到涉及协调、知识共享和共同目标的合作场景。在各层级中,区分了基于上下文的推理(通过结构化编排扩展测试时交互)与训练后推理(通过强化学习和监督微调优化行为)。我们进一步回顾了现实世界应用与基准中的代表性代理式推理框架,涵盖科学、机器人、医疗、自主研究和数学等领域。本综述将代理式推理方法整合为统一路线图,连接思维与行动,并指出未来挑战,包括个性化、长周期交互、世界建模、可扩展多智能体训练及真实部署治理。

原文摘要 · Abstract (English)

Reasoning is a fundamental cognitive process underlying inference, problem-solving, and decision-making. While large language models (LLMs) demonstrate strong reasoning capabilities in closed-world settings, they struggle in open-ended and dynamic environments. Agentic reasoning marks a paradigm shift by reframing LLMs as autonomous agents that plan, act, and learn through continual interaction. In this survey, we organize agentic reasoning along three complementary dimensions. First, we characterize environmental dynamics through three layers: foundational agentic reasoning, which establishes core single-agent capabilities including planning, tool use, and search in stable environments; self-evolving agentic reasoning, which studies how agents refine these capabilities through feedback, memory, and adaptation; and collective multi-agent reasoning, which extends intelligence to collaborative settings involving coordination, knowledge sharing, and shared goals. Across these layers, we distinguish in-context reasoning, which scales test-time interaction through structured orchestration, from post-training reasoning, which optimizes behaviors via reinforcement learning and supervised fine-tuning. We further review representative agentic reasoning frameworks across real-world applications and benchmarks, including science, robotics, healthcare, autonomous research, and mathematics. This survey synthesizes agentic reasoning methods into a unified roadmap bridging thought and action, and outlines open challenges and future directions, including personalization, long-horizon interaction, world modeling, scalable multi-agent training, and governance for real-world deployment.

智能体推理多智能体大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。