让AI学数学更聪明:按模型能力定制学习路径,难题也能变易题。
Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning
- 根据模型实际水平动态调整题目难度,不靠固定标准。
- 在5个数学推理数据集上表现显著优于传统训练方法。
- 适合需要高效提升模型推理能力的研究者与开发者。
大型语言模型在各类推理任务中表现出色,但后训练阶段存在样本利用效率低、难度样本处理僵化的问题。为此,我们提出定制化课程学习(CCL)框架,包含两项核心创新:首先,引入模型自适应难度定义,根据每个模型的个体能力定制课程数据集,而非依赖预设难度指标;其次,提出“引导提示”机制,通过战略性提示动态降低样本难度,使原本可能损害性能的难题得以有效利用。在监督微调和强化学习上的综合实验表明,CCL在五个数学推理基准上均显著优于均匀训练方法,验证了其在提升样本利用率和模型性能方面的有效性,且适用于两种范式。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have achieved remarkable performance across various reasoning tasks, yet post-training is constrained by inefficient sample utilization and inflexible difficulty samples processing. To address these limitations, we propose Customized Curriculum Learning (CCL), a novel framework with two key innovations. First, we introduce model-adaptive difficulty definition that customizes curriculum datasets based on each model's individual capabilities rather than using predefined difficulty metrics. Second, we develop "Guided Prompting," which dynamically reduces sample difficulty through strategic hints, enabling effective utilization of challenging samples that would otherwise degrade performance. Comprehensive experiments on supervised fine-tuning and reinforcement learning demonstrate that CCL significantly outperforms uniform training approaches across five mathematical reasoning benchmarks, confirming its effectiveness across both paradigms in enhancing sample utilization and model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。