arXiv:2603.00578cs.AIcs.CL2026-03被引 2

让大模型学会高效推理,减少冗余思考步骤。

Draft-Thinking: Learning Efficient Reasoning in Long Chain-of-Thought LLMs

  • 先学精简版推理路径,只保留关键步骤。
  • 在MATH500上推理成本降82.6%,性能仅降2.6%。
  • 适合追求高效推理的AI系统开发者。

长链式思维(CoT)已成为提升大推理模型(LRMs)推理能力的主流范式;然而,性能提升往往伴随推理资源消耗的显著增加。近期研究发现,现有CoT范式易引发系统性过度思考,将推理能力与成本不必要地耦合。多数先前方法通过事后技术如标记压缩、截断或长度惩罚来减少标记使用,但未直接解决推理的核心机制。我们提出Draft-Thinking,引导模型首先学习一种仅保留关键推理步骤的简洁‘草稿式’推理结构。通过渐进式课程学习,模型在规模扩展过程中稳定内化这一高效推理模式。此外,Draft-Thinking引入自适应提示,使推理深度可由模型灵活选择。大量实验表明,Draft-Thinking显著降低推理预算,同时基本保持推理性能:例如在MATH500上,推理预算减少82.6%,性能仅下降2.6%。

原文摘要 · Abstract (English)

Long chain-of-thought~(CoT) has become a dominant paradigm for enhancing the reasoning capability of large reasoning models~(LRMs); however, the performance gains often come with a substantial increase in reasoning budget. Recent studies show that existing CoT paradigms tend to induce systematic overthinking, unnecessarily coupling reasoning capability with reasoning cost. Most prior approaches reduce token usage through post hoc techniques such as token compression, truncation, or length penalties, without explicitly addressing the core mechanisms of reasoning. We propose \textbf{Draft-Thinking}, which guides models to first learn a concise \textit{draft-style} reasoning structure that retains only the critical reasoning steps. Through a \textit{progressive curriculum learning}, the model stably internalizes this efficient reasoning pattern as its capability scales. Moreover, Draft-Thinking introduces adaptive prompting, which elevates reasoning depth to a flexible, model-selectable behavior. Extensive experiments demonstrate that Draft-Thinking substantially reduces reasoning budget while largely preserving reasoning performance; for example, on MATH500, it achieves an 82.6\% reduction in reasoning budget at the cost of only a 2.6\% performance drop.

推理优化链式思维大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。