arXiv:2603.22446cs.CLcs.AI2026-03被引 23

揭示强化学习提升大模型推理的细微机制,仅少数关键词改变决定性能飞跃。

Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs

  • 通过交叉采样实验定位影响推理的关键词
  • 仅少量词替换即可恢复强化学习性能优势
  • 适合关注模型优化机理的研究者与工程师

强化学习结合可验证奖励(RLVR)显著提升了大语言模型(LLM)的推理能力,但其词级作用机制仍不清晰。本文系统研究了基线模型与强化学习模型间的分布变化,涵盖三方面:(1)词级分布偏移特征分析;(2)通过交叉采样干预评估词级偏移对序列推理的影响;(3)词级偏移的精细机制。结果表明,强化学习微调引发高度稀疏且精准的分布变化,仅极小部分词的分布发生显著偏离。我们通过词熵、位置集中度和概率质量重分配等指标刻画这些偏移的结构与演化。交叉采样实验显示,仅将少量强化学习采样的词插入基线生成中,即可逐步恢复强化学习性能;而向强化学习生成中注入少量基线词,则使性能骤降至基线水平,表明仅有少数词级决策直接贡献于性能提升。最后,我们探索了加权优势信号的变体,发现其可超越基线表现。整体结果揭示了RLVR诱导的分布变化本质,并为理解其作为靶向优化过程提供了细粒度视角。

原文摘要 · Abstract (English)

Reinforcement learning with verifiable rewards (RLVR) has significantly improved reasoning in large language models (LLMs), yet the token-level mechanisms underlying these improvements remain unclear. We present a systematic empirical study of RLVR's distributional effects organized around three main analyses: (1) token-level characterization of distributional shifts between base and RL models, (2) the impact of token-level distributional shifts on sequence-level reasoning performance through cross-sampling interventions, and (3) fine-grained mechanics of these shifts at the token level. We find that RL fine-tuning induces highly sparse and targeted changes, with only a small fraction of token distributions exhibiting meaningful divergence between the base and RL policies. We further characterize the structure and evolution of these shifts through analyses of token entropy, positional concentration, and reallocation of probability mass. To assess the functional importance of these sparse changes, we conduct cross-sampling experiments that selectively swap token choices between the base and RL models with varying intervention budgets. We show that inserting only a small fraction of RL-sampled tokens into base generations progressively recovers RL performance gains, while injecting a similarly small number of base token choices into otherwise RL-generated sequences collapses performance to base levels, isolating a small set of token-level decisions directly responsible for RLVR's performance gains. Finally, we explore divergence-weighted variants of the advantage signal as a diagnostic intervention, finding that they can yield improvements over baselines. Together, our results shed light on the distributional changes induced by RLVR and provide a fine-grained, token-level lens for understanding RLVR fine-tuning as a targeted refinement process.

强化学习大模型推理分布偏移词级分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。