arXiv:2603.27226cs.CL2026-03被引 1

实验证明,按难易排序训练对逻辑推理无明显提升

Rethinking Easy-to-Hard: Limits of Curriculum Learning in Post-Training for Deductive Reasoning

  • 用合成算术与逻辑题测试不同难度顺序的训练效果
  • 无论模型类型或训练方法,难易排序均未提升准确率或缩短回答长度
  • 适用于需要组合推理能力的模型优化,质疑传统教学法在大模型中的适用性

课程学习(CL)基于由易到难逐步学习可促进泛化这一直觉,被广泛用于大语言模型(LLM)的预训练与后训练。该直觉在组合推理任务中尤为吸引人,因复杂问题由基本推理规则构成;然而,其实际效果仍缺乏系统研究。本文针对后训练阶段的课程学习开展系统性实证研究,使用合成算术与逻辑基准测试,以推理复杂度而非表面特征定义难度。结果出人意料:在多个模型家族与课程调度下,基于难度的样本排序在准确率和响应长度上均未显著优于随机采样。该结论在监督微调(SFT)与强化学习(RL)方法中一致成立。研究表明,在演绎推理背景下,训练样本的具体排序对实现组合泛化几乎无影响,挑战了课程学习在后训练中的实际效用。

原文摘要 · Abstract (English)

Curriculum learning (CL), motivated by the intuition that learning in increasing order of difficulty should ease generalization, is commonly adopted both in pre-training and post-training of large language models (LLMs). The intuition of CL is particularly compelling for compositional reasoning, where complex problems are built from elementary inference rules; however, the actual impact of CL on such tasks remains largely underexplored. We present a systematic empirical study of CL for post-training of LLMs, using synthetic arithmetic and logical benchmarks where difficulty is characterized by reasoning complexity rather than surface-level proxies. Surprisingly, across multiple model families and curriculum schedules, we find no robust advantage in difficulty-based sequencing over standard random sampling in either accuracy or response length. These findings persist across both supervised fine-tuning (SFT) and reinforcement learning (RL) methods. Our study suggests that, in the context of deductive reasoning, the specific ordering of training examples plays a negligible role in achieving compositional generalization, challenging the practical utility of curriculum-based post-training.

课程学习逻辑推理大模型训练实证研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。