用解题数据替代通用数学文本,能显著提升大模型的数学推理能力。
Advancing Mathematical Reasoning in Language Models: The Impact of Problem-Solving Data, Data Synthesis Methods, and Training Stages
- 用解题数据替代通用数学语料进行持续预训练
- 解题数据使模型数学能力大幅提升,优于通用语料
- 解题数据在预训练阶段比微调阶段更有效
数学推理仍是大语言模型(LLM)的难点,现有数学专用模型如LLEMMA、DeepSeekMath、Qwen2-Math等通常采用两阶段训练:先用数学相关语料预训练,再用题目数据进行监督微调(SFT)。尽管如此,持续预训练(CPT)带来的数学推理提升仍远低于SFT。本研究探索预训练阶段的替代策略,聚焦于使用解题数据而非通用数学语料。我们提出三个核心问题:(1)解题数据能否在CPT中更有效地提升模型数学能力?(2)同源合成数据是否同样有效,哪些合成方法最高效?(3)相同解题数据在CPT与SFT阶段的能力差异及成因?结果表明,解题数据显著优于通用数学语料;其中,导师增强合成法表现最佳。此外,虽然SFT有助于指令遵循,但其对复杂解题任务的学习能力弱于同数据下的CPT。这些发现为优化大模型数学推理能力提供了指导,并催生了我们的数学基础模型MathGPT-8B。
原文摘要 · Abstract (English)
Mathematical reasoning remains a challenging area for large language models (LLMs), prompting the development of math-specific LLMs such as LLEMMA, DeepSeekMath, and Qwen2-Math, among others. These models typically follow a two-stage training paradigm: pre-training with math-related corpora and post-training with problem datasets for supervised fine-tuning (SFT). Despite these efforts, the improvements in mathematical reasoning achieved through continued pre-training (CPT) are often less significant compared to those obtained via SFT. This study addresses this discrepancy by exploring alternative strategies during the pre-training phase, focusing on the use of problem-solving data over general mathematical corpora. We investigate three primary research questions: (1) Can problem-solving data enhance the model's mathematical reasoning capabilities more effectively than general mathematical corpora during CPT? (2) Are synthetic data from the same source equally effective, and which synthesis methods are most efficient? (3) How do the capabilities developed from the same problem-solving data differ between the CPT and SFT stages, and what factors contribute to these differences? Our findings indicate that problem-solving data significantly enhances the model's mathematical capabilities compared to general mathematical corpora. We also identify effective data synthesis methods, demonstrating that the tutorship amplification synthesis method achieves the best performance. Furthermore, while SFT facilitates instruction-following abilities, it underperforms compared to CPT with the same data, which can be partially attributed to its poor learning capacity for more challenging problem-solving data. These insights provide valuable guidance for optimizing the mathematical reasoning capabilities of LLMs, culminating in our development of a powerful mathematical base model called MathGPT-8B.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。