arXiv:2605.09253cs.CLcs.AI2026-05被引 11

发现训练中顽固高损的'石头令牌',它们无效却消耗大量优化资源。

Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation

论文配图:Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation
图 1 · 摘自论文原文
  • 识别出持续高损失的'石头令牌',占生成文本18%以上
  • 这些令牌虽贡献大梯度但无法改进,对性能无实质帮助
  • 跳过此类令牌可显著加速对齐,适合大规模模型蒸馏场景

尽管近期强化学习中可验证奖励(RLVR)研究发现少数关键令牌显著推动推理提升,但对策略内蒸馏(OPD)的令牌级理解仍不充分。本文研究高损失令牌——根据每令牌KL目标,应随训练收敛而减少,但实证显示即使训练趋于饱和,仍存在大量持续高损失的令牌,我们称之为‘石头令牌’,可占生成输出的18%。分析揭示两大悖论:其一,尽管石头令牌高频出现且贡献巨大梯度范数,但其自身在训练中停滞不前;其二,通过因果干预发现,这些令牌对模型实际推理性能几乎无贡献。这表明大量优化资源被浪费于学生模型无法或无需内化的结构与语篇残差。通过解构此机制,我们证明有策略地绕过这些‘绊脚石’可显著精简对齐过程,挑战了统一令牌权重的必要性,为大规模模型蒸馏提供了更高效范式。

原文摘要 · Abstract (English)

While recent work in Reinforcement Learning with Verifiable Rewards (RLVR) has shown that a small subset of critical tokens disproportionately drives reasoning gains, an analogous token-level understanding of On-Policy Distillation (OPD) remains largely unexplored. In this work, we investigate high-loss tokens, a token type that--as the most direct signal of student-teacher mismatch under OPD's per-token KL objective--should progressively diminish as training converges according to existing studies; however, our empirical analysis shows otherwise. Even after OPD training reaches apparent saturation, a substantial subset of tokens continues to exhibit persistently high loss; these tokens, which we term Rock Tokens, can account for up to 18\% of the tokens in generated outputs. Our investigation reveals two startling paradoxes. First, despite their high occurrence frequency providing a disproportionately large share of total gradient norms, Rock Tokens themselves remain stagnant throughout training, resisting teacher-driven corrections. Second, through causal intervention, we find that these tokens provide negligible functional contribution to the model's actual reasoning performance. These findings suggest that a vast amount of optimization bandwidth is spent on structural and discourse residuals that the student model cannot or need not internalize. By deconstructing these dynamics, we demonstrate that strategically bypassing these ``stumbling blocks'' can significantly streamline the alignment process, challenging the necessity of uniform token weighting and offering a more efficient paradigm for large-scale model distillation.

强化学习模型蒸馏令牌分析优化效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。