构建750万条数学推理数据,支持工具调用,提升模型长文本推理能力
Nemotron-Math: Efficient Long-Context Distillation of Mathematical Reasoning from Multi-Mode Supervision
- 融合竞赛题与真实问题,生成三类风格的数学解题过程
- 在AIME 2024/2025上实现100%准确率,长上下文训练提速2-3倍
- 适合需要强数学推理和长序列处理能力的研究者使用
高质量数学推理训练需涵盖多样推理方式、长流程追踪及有效工具集成,现有数据集在此方面能力有限。我们利用gpt-oss-120b的多模态生成能力,构建了Nemotron-Math——一个大规模数学推理数据集,包含750万条解题轨迹,覆盖高、中、低三类推理模式,每种均提供带与不带Python工具调用(TIR)的版本。数据集整合8.5万道精选AoPS题目与26.2万条社区贡献的StackExchange-Math问题,结合结构化竞赛题与多样化现实数学查询。通过受控评估验证数据质量:Nemotron-Math在匹配的AoPS问题上持续优于原始OpenMathReasoning;引入StackExchange-Math显著提升对HLE-Math的鲁棒性与泛化能力,同时保持竞赛基准上的高准确率。为支持高效长上下文训练,我们设计顺序分桶策略,使128K上下文长度微调加速2–3倍且精度损失极小。整体上,Nemotron-Math实现了顶尖性能,包括在AIME 2024与2025上使用Python TIR达到100% maj@16准确率。
原文摘要 · Abstract (English)
High-quality mathematical reasoning supervision requires diverse reasoning styles, long-form traces, and effective tool integration, capabilities that existing datasets provide only in limited form. Leveraging the multi-mode generation ability of gpt-oss-120b, we introduce Nemotron-Math, a large-scale mathematical reasoning dataset containing 7.5M solution traces across high, medium, and low reasoning modes, each available both with and without Python tool-integrated reasoning (TIR). The dataset integrates 85K curated AoPS problems with 262K community-sourced StackExchange-Math problems, combining structured competition tasks with diverse real-world mathematical queries. We conduct controlled evaluations to assess the dataset quality. Nemotron-Math consistently outperforms the original OpenMathReasoning on matched AoPS problems. Incorporating StackExchange-Math substantially improves robustness and generalization, especially on HLE-Math, while preserving accuracy on math competition benchmarks. To support efficient long-context training, we develop a sequential bucketed strategy that accelerates 128K context-length fine-tuning by 2--3$\times$ without significant accuracy loss. Overall, Nemotron-Math enables state-of-the-art performance, including 100\% maj@16 accuracy on AIME 2024 and 2025 with Python TIR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。