arXiv:2507.02076cs.AIcs.LG2025-07综述被引 30

让大模型按需思考,简单问题少算,复杂问题多算。

Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs

  • 分两层控制推理计算:固定预算和动态调整。
  • 实测显示不同任务下计算量与推理效果存在显著权衡。
  • 适合关注模型效率与响应速度的开发者和研究者。

大语言模型(LLMs)已发展为能处理多种任务的通用智能体,但推理效率仍不足:无论任务难易,均使用固定的推理计算量,常对简单问题过度计算,对难题又计算不足。本综述系统梳理了高效测试时计算(TTC)策略,旨在提升LLM推理的计算效率。我们提出双层分类体系,区分L1可控性(固定计算预算下的方法)与L2自适应性(根据输入难度或模型置信度动态调整推理的方法)。在多个数据集上评测主流专有大型语言模型,揭示了推理性能与令牌使用量之间的关键权衡。相比以往关于高效推理的综述,本工作更强调TTC方法在实际应用中的可控性、适应性与可扩展性。最后,探讨了混合思维模型等新兴趋势,并指出了未来实现更高效、鲁棒且符合用户约束的LLM的关键挑战。

原文摘要 · Abstract (English)

Large language models (LLMs) have rapidly progressed into general-purpose agents capable of solving a broad spectrum of tasks. However, current models remain inefficient at reasoning: they apply fixed inference-time compute regardless of task complexity, often overthinking simple problems while underthinking hard ones. This survey presents a comprehensive review of efficient test-time compute (TTC) strategies, which aim to improve the computational efficiency of LLM reasoning. We introduce a two-tiered taxonomy that distinguishes between L1-controllability, methods that operate under fixed compute budgets, and L2-adaptiveness, methods that dynamically scale inference based on input difficulty or model confidence. We benchmark leading proprietary LLMs across diverse datasets, highlighting critical trade-offs between reasoning performance and token usage. Compared to prior surveys on efficient reasoning, our review emphasizes the practical control, adaptability, and scalability of TTC methods. Finally, we discuss emerging trends such as hybrid thinking models and identify key challenges for future work towards making LLMs more computationally efficient, robust, and responsive to user constraints.

大模型推理计算效率自适应计算测试时优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。