arXiv:2503.24377cs.CLcs.AI2025-03综述被引 27

如何让大模型推理更快更省,这篇综述给出了系统性解决方案。

Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models

  • 分析大模型推理效率低的根源与不同思考模式的行为差异
  • 提出平衡性能与计算成本的推理经济新范式
  • 适合关注模型效率、推理优化的研究者与工程师

大语言模型(LLMs)在复杂推理任务上的进步,使其从快速直觉思维(System 1)转向缓慢深入的思考(System 2)。尽管System 2提升任务准确率,但其高计算开销源于慢速思维及冗余推理行为。而System 1虽高效却表现不佳。因此,必须在性能收益与计算预算间取得平衡,催生了“推理经济”概念。本文全面分析了后训练与推理阶段的推理经济问题,涵盖:推理低效的成因、不同推理模式的行为特征,以及实现推理经济的潜在方案。通过提供可操作洞察并指出开放挑战,旨在为提升大模型推理经济性提供指导,并附带公开仓库持续追踪该领域进展。

原文摘要 · Abstract (English)

Recent advancements in Large Language Models (LLMs) have significantly enhanced their ability to perform complex reasoning tasks, transitioning from fast and intuitive thinking (System 1) to slow and deep reasoning (System 2). While System 2 reasoning improves task accuracy, it often incurs substantial computational costs due to its slow thinking nature and inefficient or unnecessary reasoning behaviors. In contrast, System 1 reasoning is computationally efficient but leads to suboptimal performance. Consequently, it is critical to balance the trade-off between performance (benefits) and computational costs (budgets), giving rise to the concept of reasoning economy. In this survey, we provide a comprehensive analysis of reasoning economy in both the post-training and test-time inference stages of LLMs, encompassing i) the cause of reasoning inefficiency, ii) behavior analysis of different reasoning patterns, and iii) potential solutions to achieve reasoning economy. By offering actionable insights and highlighting open challenges, we aim to shed light on strategies for improving the reasoning economy of LLMs, thereby serving as a valuable resource for advancing research in this evolving area. We also provide a public repository to continually track developments in this fast-evolving field.

推理效率大模型系统2算力优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。