arXiv:2609.08341cs.LG2026-09

通过约束梯度符号与总变差,提升强化学习中策略蒸馏的稳定性

TV-Regulated OPD: Direction Matters in On-Policy Distillation

论文配图:TV-Regulated OPD: Direction Matters in On-Policy Distillation
图 1 · 摘自论文原文
  • 仅保留词级别优势的符号信息即可达到标准蒸馏效果
  • 引入总变差正则使优势函数更平滑,训练方差降低37%以上
  • 适合追求稳定训练过程的LLM微调场景

在大语言模型后训练阶段,基于策略的蒸馏(OPD)用于将领域专家知识传递给学生模型。然而主流方法中的监督信号存在高方差和噪声,导致训练不稳定。本文系统研究了影响性能的关键因素及不稳定性机制,发现仅保留词级优势的符号信息即可达到与标准OPD相当的性能。同时,更平滑且有界的优点可显著稳定训练过程。为此,我们利用总变差(TV)对优势函数进行约束,提出鲁棒的TV-OPD方法。得益于有界且被抑制的优势,TV-OPD展现出稳定的训练动态与一致的后期性能。在多种设置下的全面实验表明,TV-OPD在后期训练阶段始终实现更优性能与更低方差。

原文摘要 · Abstract (English)

On-Policy Distillation (OPD) facilitates the transfer of knowledge from domain expert to student in the post-training phase of Large Language Models (LLMs). However, the supervision signals in mainstream OPD methods suffer from high variance and noise which is generally instable during training. In this work, we systematically investigated what really matters to the performance and the fundamental mechanisms behind the instability during training. We found that retaining only the sign of token-level advantages is sufficient to achieve the performance comparable to standard OPD. Meanwhile, smoother and bounded advantages can stabilize the training process without sacrificing its performance. These motivated us to shape the advantages using the Total Variation (TV) and propose a robust TV regulated On-Policy Distillation (TV-OPD) method. Benefiting from the bounded and diminished advantages, TV-OPD exhibits stable training dynamics and steady late-stage performance. We conducted comprehensive experiments and found that, across various settings, TV-OPD consistently achieved better performance and lower variance in the late-stage of training.

强化学习知识蒸馏大模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。