arXiv:2509.21124cs.AIcs.CL2025-09被引 4

通过挖掘高价值推理模式,用少量数据显著提升大模型的解题能力。

Expanding Reasoning Potential in Foundation Model by Learning Diverse Chains of Thought Patterns

  • 抽象出通用且可推广的推理原子模式,构建核心参考集。
  • 仅用100亿token高质量推理数据,使模型在AIME上提升9.58%。
  • 适合想高效训练数学推理模型的研究者和开发者。

大型推理模型在复杂数学推理任务上的进展主要依赖强化学习。在训练中期引入长链式思维(CoT)数据已被证明能显著提升推理深度。然而,现有方法常盲目使用CoT数据,未解决何种数据最有效的问题。本文首次将基础模型的推理潜力定义为正确回答问题所需独立尝试次数的倒数,该指标与最终性能强相关。为此,我们提出利用富含高价值推理模式的数据来扩展推理潜力。具体而言,从CoT序列中抽象出具有普遍性和归纳能力的原子推理模式,构建核心参考集。进一步设计双粒度算法,结合推理模式链与词元熵,从数据池中高效筛选与核心集匹配的高价值CoT数据(CoTP),从而让模型有效掌握推理技能。仅使用100亿token的CoTP数据,即能使850亿参数的MoE模型在挑战性AIME 2024和2025测试集上提升9.58%,并将下游强化学习性能上限提高7.81%。

原文摘要 · Abstract (English)

Recent progress in large reasoning models for challenging mathematical reasoning has been driven by reinforcement learning (RL). Incorporating long chain-of-thought (CoT) data during mid-training has also been shown to substantially improve reasoning depth. However, current approaches often utilize CoT data indiscriminately, leaving open the critical question of which data types most effectively enhance model reasoning capabilities. In this paper, we define the foundation model's reasoning potential for the first time as the inverse of the number of independent attempts required to correctly answer the question, which is strongly correlated with the final model performance. We then propose utilizing diverse data enriched with high-value reasoning patterns to expand the reasoning potential. Specifically, we abstract atomic reasoning patterns from CoT sequences, characterized by commonality and inductive capabilities, and use them to construct a core reference set enriched with valuable reasoning patterns. Furthermore, we propose a dual-granularity algorithm involving chains of reasoning patterns and token entropy, efficiently selecting high-value CoT data (CoTP) from the data pool that aligns with the core set, thereby training models to master reasoning effectively. Only 10B-token CoTP data enables the 85A6B Mixture-of-Experts (MoE) model to improve by 9.58% on the challenging AIME 2024 and 2025, and to raise the upper bound of downstream RL performance by 7.81%.

推理增强数学推理MoE模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。