arXiv:2607.13399cs.CLcs.LG2026-07被引 5

揭示大模型对齐中在线蒸馏的机制与陷阱,提出有效调控信号质量的方法。

Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations

论文配图:Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations
图 1 · 摘自论文原文
  • 将在线蒸馏视为探索催化剂,依赖高质量逐标记引导信号。
  • 发现教师-学生分布差异和长度依赖捷径会导致探索失效。
  • 通过优势截断和对数压缩轻量调节信号,提升蒸馏稳定性与效果。

在线蒸馏(OPD)已成为大语言模型后训练的核心范式,但其训练动态仍不清晰。本文系统研究了OPD的作用、病理及调控机制。首先明确其角色为探索催化剂:通过密集的逐标记级指导,引导学生走向正确推理路径,而不提升能力上限。实验表明,提示多样性比单题采样数量更重要,且OPD效果完全取决于引导信号质量。该依赖暴露两大病理:当教师与学生分布差距过大时,出现学生-教师错配,导致引导信号偏离任务正确性;当聚合的逐标记目标产生长度依赖捷径时,引发长度滥用,使学生通过截断或冗余填充“投机”奖励,而非学习真实推理策略。为抑制这些病理,本文探索轻量级信号调控方法——优势截断与对数尺度压缩,确保探索由忠实信号引导。在七个基准测试上的实验表明,这些调控可缓解长度滥用,实现稳定超越现有OPD变体与RLVR基线,证实良好调控的信号质量才是驱动成功探索的关键,而非单纯教师规模。

原文摘要 · Abstract (English)

On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly understood. We present a systematic study examining the role, pathologies, and regulations of OPD. We first clarify the role of OPD as an exploration catalyst: it steers the student toward correct reasoning paths via dense token-level guidance, without expanding capability ceiling. We confirm this by showing that prompt diversity matters more than per-problem sampling numbers, and critically, that the effectiveness of OPD hinges entirely on the quality of its guiding signal. This dependency exposes two pathologies that derail exploration. The Student-Teacher Mismatch occurs when a large teacher-student distributional gap causes the guiding signal to misalign with task correctness, steering exploration in counterproductive directions. Length Exploitation arises when the aggregated token-level objective creates length-dependent shortcuts, allowing the student to game the reward landscape through response truncation or redundant padding, exploring degenerate length modes rather than reasoning strategies. To tame these pathologies, we investigate lightweight signal regulations: advantage clipping and log-scale compression, ensuring exploration is guided by faithful signals. Experiments across seven benchmarks demonstrate that these regulations alleviate length exploitation and enable effective distillation, stably surpassing OPD variants and RLVR baselines, thereby confirming that well-regulated signal quality, rather than mere teacher scale, governs successful exploration in OPD.

大模型训练知识蒸馏强化学习信号调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。