arXiv:2512.21625cs.CL2025-12被引 17

通过调整正负样本优势值,提升大模型强化学习的推理能力

Rethinking Sample Polarity in Reinforcement Learning with Verifiable Rewards

  • 区分正负样本,分别优化正确路径与探索新思路
  • 在5个推理基准上显著提升模型表现
  • 适合研究大模型强化学习与推理优化的读者

大型推理模型(LRMs)通常通过可验证奖励的强化学习(RLVR)训练以增强推理能力。在此范式中,策略利用自生成的正样本和负样本进行更新,对应不同样本极性。本文系统研究了样本极性对RLVR训练动态的影响:正样本强化已有正确推理模式,负样本促进探索新推理路径。进一步分析了在样本级和标记级调整正负样本优势值的影响,提出一种自适应、非对称的标记级优势塑造方法A3PO,更精准地将优势信号分配给不同极性的关键标记。在五个推理基准上的实验验证了该方法的有效性。

原文摘要 · Abstract (English)

Large reasoning models (LRMs) are typically trained using reinforcement learning with verifiable reward (RLVR) to enhance their reasoning abilities. In this paradigm, policies are updated using both positive and negative self-generated rollouts, which correspond to distinct sample polarities. In this paper, we provide a systematic investigation into how these sample polarities affect RLVR training dynamics and behaviors. We find that positive samples sharpen existing correct reasoning patterns, while negative samples encourage exploration of new reasoning paths. We further explore how adjusting the advantage values of positive and negative samples at both the sample level and the token level affects RLVR training. Based on these insights, we propose an Adaptive and Asymmetric token-level Advantage shaping method for Policy Optimization, namely A3PO, that more precisely allocates advantage signals to key tokens across different polarities. Experiments across five reasoning benchmarks demonstrate the effectiveness of our approach.

强化学习推理模型优势塑造样本极性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。