arXiv:2602.08354cs.AI2026-02被引 23

让大模型自己决定何时停止思考,提升推理效率与准确率

Does Your Reasoning Model Implicitly Know When to Stop Thinking?

  • 设计新采样策略SAGE,让模型自主判断何时停止推理
  • 在多个数学基准上实现更高准确率与更快推理速度
  • 适合追求高效推理的AI系统开发者使用

大型推理模型(LRMs)通过长思维链(CoTs)显著提升了复杂推理任务的表现,但常导致大量冗余,降低计算效率并延缓实时响应。研究发现,更长的推理链往往与正确性无关,甚至损害准确率。我们深入分析后意外发现:LRMs实际上隐含知道何时该停止思考,只是当前采样方式掩盖了这一能力。为此,我们提出SAGE(自感知引导的高效推理)采样策略,释放模型的高效推理潜力。进一步地,将SAGE作为混合采样方式融入基于群体的强化学习(SAGE-RL),使SAGE-RL能有效整合高效推理模式,在标准pass@1推理中显著提升多个挑战性数学基准上的推理准确率与效率。

原文摘要 · Abstract (English)

Recent advancements in large reasoning models (LRMs) have greatly improved their capabilities on complex reasoning tasks through Long Chains of Thought (CoTs). However, this approach often results in substantial redundancy, impairing computational efficiency and causing significant delays in real-time applications. Recent studies show that longer reasoning chains are frequently uncorrelated with correctness and can even be detrimental to accuracy. In a further in-depth analysis of this phenomenon, we surprisingly uncover and empirically verify that LRMs implicitly know the appropriate time to stop thinking, while this capability is obscured by current sampling paradigms. Motivated by this, we introduce SAGE (Self-Aware Guided Efficient Reasoning), a novel sampling paradigm that unleashes this efficient reasoning potential. Furthermore, integrating SAGE as mixed sampling into group-based reinforcement learning (SAGE-RL) enables SAGE-RL to effectively incorporate SAGE-discovered efficient reasoning patterns into standard pass@1 inference, markedly enhancing both the reasoning accuracy and efficiency of LRMs across multiple challenging mathematical benchmarks.

推理优化思维链高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。