arXiv:2605.25381cs.LG2026-05

通过动态调整学习时机,提升大模型强化学习的稳定性和效率。

Not only where, But when: Temporal Scheduling for RLVR

论文配图:Not only where, But when: Temporal Scheduling for RLVR
图 1 · 摘自论文原文
  • 引入时间调度机制,随训练进程动态调整奖励分配策略。
  • 在数学与通用推理任务上,显著提升学习稳定性和收敛速度。
  • 适合追求高效、稳定微调大模型的研究者和工程师。

基于可验证奖励的强化学习(RLVR)已成为大语言模型后训练的核心技术。现有方法虽通过逐标记优势重加权或选择性优化来处理轨迹中异质策略行为,但奖励分配标准在整个训练过程中保持不变,限制了策略的韧性进化。本文提出,学习信号的「何时」调度与「何处」分配同样重要,引入时间维度,动态调整信用分配准则。实验发现,初期聚焦特定行为标记,逐步转向泛化优化,能带来更稳定高效的训练动态。此外,简单使用轨迹分位数即可有效区分策略行为,配合时间调度表现优异。分析表明,标准优化会显著降低策略熵以兼顾异质行为,而时间调度则促进更健康的策略演化。在数学与通用推理基准上的实验均显示一致性能提升,证明时间调度是值得探索的重要优化维度。

原文摘要 · Abstract (English)

Reinforcement learning with verifiable rewards (RLVR) has become a core technique for post-training of Large Language Models (LLMs). While policy optimization is driven by all sampled tokens under a globally broadcast scalar reward, the heterogeneous policy behaviors exhibited along trajectories are largely overlooked without differentiation. Existing works address this by credit allocation, including token-level advantage reweighting, and selective token optimization, however, the allocation criterion are principally stagnant throughout training, limiting resilient policy evolution. In this work, we argue that \textit{when} learning signals are scheduled can be as important as \textit{where} they are allocated across tokens, and introduce the temporal dimension that scheduling the credit allocation criteria over the course of RLVR optimization. We find that prioritizing targeted tokens emphasized with specific policy behaviors, and gradually attenuating toward general optimization leads to more stable and efficient learning dynamics. Furthermore, we show that simple trajectory percentiles provide a natural perspective for distinguishing policy behaviors, and works effectively with temporal scheduling. Our analysis reveals that standard optimization substantially sacrifices policy entropy when simultaneously accommodating heterogeneous behaviors, whereas temporal scheduling yields healthier policy evolution dynamics. Experiments across mathematical and general reasoning benchmarks demonstrate consistent improvements, suggesting that temporal scheduling constitutes a promising optimization dimension.

强化学习大模型微调奖励调度策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。