arXiv:2603.22117cs.LGcs.AI2026-03被引 12

发现强化学习更新方向比幅度更能揭示大模型推理提升机制

On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation

  • 用符号化的概率差Δlog p捕捉每令牌更新方向
  • 方向指标比幅度指标更精准定位关键推理更新
  • 可应用于推理时外推与训练时加权,无需额外训练

基于可验证奖励的强化学习(RLVR)显著提升了大语言模型的推理能力。现有研究多关注更新的幅度(如发散或熵),却忽视了更新方向的重要性。本文提出用基线模型与最终模型之间的符号化令牌级对数概率差Δlog p来表征更新方向。通过统计分析与替换干预实验,证明Δlog p比幅度指标更有效地识别出稀疏但对推理至关重要的更新。基于此,本文提出两项应用:(1) 推理时沿Δlog p方向外推策略,无需再训练即可提升推理准确率;(2) 训练时对低概率(对应高Δlog p)令牌进行重加权,可在多个模型与基准上提升推理表现。本工作确立更新方向为分析和改进RLVR的核心原则。

原文摘要 · Abstract (English)

Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning capabilities of large language models. While existing analyses identify that RLVR-induced changes are sparse, they primarily focus on the \textbf{magnitude} of these updates, largely overlooking their \textbf{direction}. In this work, we argue that the direction of updates is a more critical lens for understanding RLVR's effects, which can be captured by the signed, token-level log probability difference $Δ\log p$ between the base and final RLVR models. Through statistical analysis and token-replacement interventions, we demonstrate that $Δ\log p$ more effectively identifies sparse, yet reasoning-critical updates than magnitude-based metrics (\eg divergence or entropy). Building on this insight, we propose two practical applications: (1) a \textit{test-time extrapolation} method that amplifies the policy along the learned $Δ\log p$ direction to improve reasoning accuracy without further training; (2) a \textit{training-time reweighting} method that focuses learning on low-probability (corresponding to higher $Δ\log p$) tokens, which improves reasoning performance across models and benchmarks. Our work establishes the direction of change as a key principle for analyzing and improving RLVR.

强化学习大模型推理更新方向零样本外推

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。