通过可控难度的数据生成与课程学习,显著提升大模型数学推理能力。
MathMixup: Boosting LLM Mathematical Reasoning with Difficulty-Controllable Data Synthesis and Curriculum Learning
- 采用混合分解策略生成可调节难度的数学题,确保题目质量与梯度
- 在7个基准上使Qwen2.5-7B平均得分达52.6%,超越现有最优方法
- 适合需要强化数学推理的大模型训练,尤其适用于数据驱动的课程学习
在数学推理任务中,大语言模型的进步高度依赖高质量、难度分级明确的训练数据。然而,现有数据生成方法常因多样性不足且难以精确控制题目难度,难以支持如课程学习等高效训练范式。为此,我们提出MathMixup,一种通过混合与分解策略系统生成高质量、难度可控数学推理题的新范式。结合自动自检与人工筛选,确保题目语义清晰且具备良好难度梯度。基于此,我们构建了MathMixupQA数据集,并设计了利用分级题目的课程学习策略,支持与其他数据集灵活融合。实验表明,MathMixup及其课程学习策略显著提升了大模型的数学推理性能。微调后的Qwen2.5-7B在七个数学基准上平均得分为52.6%,优于此前最先进方法。结果充分验证了MathMixup在提升大模型数学推理能力及推动以数据为中心的课程学习方面的有效性与普适性。
原文摘要 · Abstract (English)
In mathematical reasoning tasks, the advancement of Large Language Models (LLMs) relies heavily on high-quality training data with clearly defined and well-graded difficulty levels. However, existing data synthesis methods often suffer from limited diversity and lack precise control over problem difficulty, making them insufficient for supporting efficient training paradigms such as curriculum learning. To address these challenges, we propose MathMixup, a novel data synthesis paradigm that systematically generates high-quality, difficulty-controllable mathematical reasoning problems through hybrid and decomposed strategies. Automated self-checking and manual screening are incorporated to ensure semantic clarity and a well-structured difficulty gradient in the synthesized data. Building on this, we construct the MathMixupQA dataset and design a curriculum learning strategy that leverages these graded problems, supporting flexible integration with other datasets. Experimental results show that MathMixup and its curriculum learning strategy significantly enhance the mathematical reasoning performance of LLMs. Fine-tuned Qwen2.5-7B achieves an average score of 52.6\% across seven mathematical benchmarks, surpassing previous state-of-the-art methods. These results fully validate the effectiveness and broad applicability of MathMixup in improving the mathematical reasoning abilities of LLMs and advancing data-centric curriculum learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。