arXiv:2608.03068cs.CLcs.AI2026-08

通过动态课程与方差感知优化,提升大模型强化学习推理能力

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning

论文配图:CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning
图 1 · 摘自论文原文
  • 引入轨迹级方差感知优势调整,更精准反馈生成过程
  • 动态课程权重匹配问题难度,避免能力与任务错配
  • 在数学推理任务中显著优于现有基线,探索更充分

强化学习已成为提升大语言模型推理能力的有效方法。然而,现有方法在生成答案轨迹的反馈精度不足,且存在问题难度漂移现象。为此,我们提出CVPO——基于课程引导的价值-方差策略优化。在响应轨迹层面,发现分词级价值方差与探索强度相关,理论分析表明该方差可约束策略更新幅度。据此设计了针对不同奖励类型的方差感知优势调整机制。在问题层面,引入动态课程加权方法,使模型在训练各阶段聚焦于与其当前能力匹配的任务。实验结果表明,该方法优于强基线VAPO,实现更高性能与更强探索能力,显著提升模型在各类数学任务中的推理准确性与鲁棒性。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has emerged as an effective method for enhancing the reasoning capabilities of large language models (LLMs). However, existing methods suffer from insufficient precision in feedback on generated answer trajectories and exhibit the phenomenon of problem difficulty drift. To address these challenges, we propose CVPO - Curriculum-guided Value-Variance Policy Optimization. At the response trajectory level, we find that token-level value-variance correlates with exploration intensity. Our theoretical analysis shows this variance bounds policy update magnitude. We then use the estimated trajectory value-variance to quantify the intrinsic randomness in generation. Based on this, we design a variance-aware advantage adjustment mechanism for different reward types. At the question level, we introduce a dynamic curriculum weighting method that adapts to question difficulty. This helps the model focus on tasks matched to its current ability during each training stage. Experimental results show our method outperforms strong value-based baselines like VAPO. It achieves better performance and stronger exploration, enabling more accurate and robust reasoning in language models across various math tasks.

强化学习大模型推理动态课程方差感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。