用动态近邻教师提升策略蒸馏稳定性,零额外开销
Trust Region Policy Distillation
- 构建动态近邻教师,降低梯度方差
- 训练更稳定,样本效率与最终性能显著提升
- 适合追求高效可靠策略蒸馏的从业者
大目标难以一蹴而就;分步实现更为明智。我们提出信任区域策略蒸馏(TOP-D),通过动态构建近邻教师,将众所周知不稳定、高方差的在线策略蒸馏(OPD)转化为稳定的训练范式。理论上,我们建立严格框架,证明了TOP-D能内在控制梯度方差。通过提供形式化的全局收敛分析及单调改进界,数学化地阐明了整体训练动态的可靠性和稳定性。实验上,TOP-D在数学推理任务中显著提升训练稳定性、样本效率和最终性能。更重要的是,TOP-D不引入任何额外计算开销,使其成为成熟OPD范式的有力替代方案。
原文摘要 · Abstract (English)
Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable, high-variance On-Policy Distillation (OPD) into a stable training paradigm by dynamically constructing a proximal teacher. Theoretically, we establish a rigorous framework demonstrating that TOP-D inherently controls gradient variance. By providing a formal global convergence analysis alongside a monotonic improvement bound, we mathematically formalize the reliability and stability of the overall training dynamics. Empirically, TOP-D dramatically enhances training stability, sample efficiency, and final performance on mathematical reasoning tasks. More importantly, TOP-D introduces zero additional computational overhead, positioning itself as a promising alternative to the well-established OPD paradigm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。