arXiv:2602.12579cs.LGcs.AI2026-02被引 7

提出无需外部验证的强化学习框架,通过模型自信度稳定训练过程。

VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction

  • 用模型自身置信度构建无验证器的课程学习机制
  • 在数学与通用推理任务上超越有无验证器的基线方法
  • 理论证明估计器渐近无偏,适合追求稳定训练的LLM研究者

基于可验证奖励的强化学习(RLVR)已成为提升大语言模型推理能力的主流范式,但其依赖外部验证器限制了可扩展性。近期研究表明,RLVR主要作用在于激发模型隐含能力,推动无验证器算法的发展。然而,在此类设置下,标准方法如组相对策略优化面临破坏性梯度方差问题,常导致训练崩溃。为此,我们提出无验证器课程强化学习(VI-CuRL),利用模型内在置信度构建独立于外部验证器的课程。通过优先选择高置信度样本,VI-CuRL有效管理偏差-方差权衡,特别降低动作与问题层面的方差。我们提供严格的理论分析,证明所提估计器保证渐近无偏性。实验表明,VI-CuRL在数学与通用推理基准上均实现训练稳定性,并持续优于有/无验证器的基线方法。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a dominant paradigm for enhancing Large Language Models (LLMs) reasoning, yet its reliance on external verifiers limits its scalability. Recent findings suggest that RLVR primarily functions by eliciting latent capabilities, motivating the development of verifier-free algorithms. However, in such settings, standard methods like Group Relative Policy Optimization face a critical challenge: destructive gradient variance that often leads to training collapse. To address this issue, we introduce Verifier-Independent Curriculum Reinforcement Learning (VI-CuRL), a framework that leverages the model's intrinsic confidence to construct a curriculum independent from external verifiers. By prioritizing high-confidence samples, VI-CuRL effectively manages the bias-variance trade-off, specifically targeting the reduction of action and problem variance. We provide a rigorous theoretical analysis, proving that our estimator guarantees asymptotic unbiasedness. Empirically, VI-CuRL promotes stability and consistently outperforms verifier-dependent/independent baselines across math and general reasoning benchmarks with/without verifiers.

强化学习大模型推理无验证器训练稳定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。