arXiv:2605.06523cs.LGcs.AI2026-05被引 2

发现强化学习推理模型存在隐式奖励过拟合,且主要集中在秩1成分。

On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR

论文配图:On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR
图 1 · 摘自论文原文
  • 通过周期性秩1替换,揭示模型在训练中对奖励的隐式过拟合现象。
  • 模型在测试集表现良好时,训练阶段奖励仍较低,说明其依赖特定秩1组件。
  • 适合关注强化学习机制、模型参数演化及持续学习的研究者。

近期研究表明,通过可验证奖励强化学习(RLVR)获得的增强推理能力主要集中于秩1成分。基于此,我们采用周期性秩1替换,发现一个反直觉现象:RLVR可能对训练数据集产生隐式奖励过拟合。具体而言,即使模型在训练过程中奖励保持相对较低,也能在测试集上取得满意表现。此外,我们刻画了三种训练特性:(1) RLVR中的有效秩1成分仅保留数学推理能力,不包含其他知识;(2) RLVR本质上是优化特定奇异谱,几乎所有线性层的奇异值分布呈现重尾特征;(3) 秩1成分对应的左奇异向量在训练中表现出更强对齐趋势,表明RLVR实质上在优化采样效率。这些发现揭示了RLVR如何塑造模型参数,为改进现有强化学习范式或实现持续学习提供了潜在洞见。

原文摘要 · Abstract (English)

Recent extensive research has demonstrated that the enhanced reasoning capabilities acquired by models through Reinforcement Learning with Verifiable Rewards (RLVR) are primarily concentrated within the rank-1 components. Predicated on this observation, we employed Periodic Rank-1 Substitution and identified a counterintuitive phenomenon: RLVR may exhibit implicit reward overfitting to the training dataset. Specifically, the model can achieve satisfactory performance on the test set even when its rewards remain relatively low during the training process. Furthermore, we characterize three distinct properties of RL training: (1) The effective rank-1 component in RLVR don't maintain other model knowledge except mathematical reasoning capability. (2) RLVR fundamentally functions by optimizing a specific singular spectrum. The distribution of singular values of almost all linear layers in RLVR-trained model behaves like heavy-tailed distribution. (3) the left singular vectors associated with rank-1 components demonstrate a stronger alignment tendency during training, which echoes the discovery that RLVR is optimizing sampling efficiency in essence. Taken together, our findings and analysis further reveal how RLVR shapes model parameters and offer potential insights for improving existing RL paradigms or other training paradigms to implement continual learning.

强化学习模型分析秩1结构持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。