arXiv:2512.19585cs.CL2025-12

增加推理时间未必提升性能,优化策略比堆算力更有效

Increasing the Thinking Budget is Not All You Need

  • 系统测试不同推理配置对性能的影响
  • 自洽和自我反思比延长推理时间更准确
  • 适合关注推理效率与成本的模型研究者

近期涌现的具备思考能力的大语言模型在多项推理基准上表现优异。早期研究开始探索推理过程长度(即思考预算)对模型性能的影响。本文系统性地考察了思考预算这一关键参数,分析其与自洽性、反思等配置的交互作用。目标是建立兼顾性能与计算成本的平衡评估框架。研究发现,单纯增加思考预算并非最有效的算力利用方式;通过自洽性和自反思等替代配置,反而能获得更高准确率。

原文摘要 · Abstract (English)

Recently, a new wave of thinking-capable Large Language Models has emerged, demonstrating exceptional capabilities across a wide range of reasoning benchmarks. Early studies have begun to explore how the amount of compute in terms of the length of the reasoning process, the so-called thinking budget, impacts model performance. In this work, we propose a systematic investigation of the thinking budget as a key parameter, examining its interaction with various configurations such as self-consistency, reflection, and others. Our goal is to provide an informative, balanced comparison framework that considers both performance outcomes and computational cost. Among our findings, we discovered that simply increasing the thinking budget is not the most effective use of compute. More accurate responses can instead be achieved through alternative configurations, such as self-consistency and self-reflection.

大模型推理思考预算自洽性效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。