arXiv:2602.14872cs.LGcs.AI2026-02被引 6

发现强化学习中隐式课程机制可自动由易到难推进长程推理训练。

On the Emergence of Implicit Curriculum in RLVR Learning Dynamics

  • 通过理论分析揭示:混合难度训练会自发形成从易到难的学习顺序。
  • 平滑难度谱下训练进入稳定接力状态,梯度信号持续推动进步。
  • 适合研究大模型推理训练机制或优化策略的学者参考。

基于可验证奖励的强化学习(RLVR)是近期大型推理模型突破的关键驱动力。然而,仅依赖最终结果的奖励如何克服长程推理障碍仍不明确。本文针对变换器在组合推理任务上的RLVR训练动态,建立理论框架。结果表明,混合难度训练会自然诱导出隐式课程:无需显式调度,简单问题先被掌握,并为更复杂问题设定学习边界,形成由易到难的学习进程。该课程有效性取决于难度谱的平滑性:当难度变化平滑时,训练进入稳定的接力阶段,较简单问题的持续梯度信号使稍难问题可解,维持在能力边缘;若难度存在突变,则出现类似‘顿悟’的相变,伴随长期停滞后才重启进展。技术上,本工作发展并适配了有限群上的傅里叶分析方法。通过受控合成实验与真实模型的RLVR运行,验证了预测机制。

原文摘要 · Abstract (English)

Reinforcement learning with verifiable rewards (RLVR) has been a main driver of recent breakthroughs in large reasoning models. Yet it remains a mystery how rewards based solely on final outcomes can help overcome the long-horizon barrier to extended reasoning. To understand this, we develop a theory of the training dynamics of RLVR for transformers on compositional reasoning tasks. Our theory shows that mixed-difficulty training naturally induces an implicit curriculum: without any explicit schedule, easier problems become learnable first and shape the frontier for harder ones, creating a learning progression from easy to hard during optimization. The effectiveness of this curriculum is governed by the smoothness of the difficulty spectrum. When the spectrum is smooth, training dynamics enter a well-behaved relay regime, in which persistent gradient signals on easier problems make slightly harder ones tractable and keep training at the edge of competence. When the spectrum contains abrupt discontinuities, training undergoes grokking-type phase transitions with prolonged plateaus before progress recurs. As a technical contribution, our analysis develops and adapts techniques from Fourier analysis on finite groups to our setting. We validate the predicted mechanisms empirically via controlled synthetic experiments and real-model RLVR runs.

强化学习推理模型隐式课程训练动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。