用教师熵动态调节学生模型模仿强度,提升小模型性能。
Rethinking Reverse KL as Adaptive Entropy Distillation
- 基于反向KL分解出教师拟合与学生熵项,无需显式正向KL。
- 理论证明最优学生分布是带温度调整的教师分布,平衡模式寻找与不确定性保留。
- 在指令遵循和数学推理任务上表现更优,适合追求高保真生成的场景。
知识蒸馏广泛用于将大语言模型的能力迁移至小型学生模型,但现有目标常难以平衡忠实模仿与鲁棒生成。现有方法多结合前向KL(FKL)与反向KL(RKL),却忽略了RKL本身具备调节学生模仿强度的机制。本文重新审视基于策略的反向Kullback-Leibler(RKL)蒸馏,将其目标分解为教师拟合项与学生熵项,无需引入显式FKL分支。理论表明,最优令牌级学生分布对应于教师分布的温度变体,其中自适应权重控制模式寻找与不确定性保留之间的权衡。基于此,提出自适应熵蒸馏(AED),利用教师熵动态校准令牌级模仿强度。在指令遵循与数学推理基准上的实验表明,AED在整体性能上表现更优,并显著提升师生分布与熵对齐效果。
原文摘要 · Abstract (English)
Knowledge distillation (KD) is widely used to transfer the capabilities of large language models (LLMs) to smaller students, but existing objectives often struggle to balance faithful imitation and robust generation. In particular, existing methods mainly combine FKL and RKL, overlooking that RKL itself provides a mechanism for adjusting the student's imitation strength. Motivated by this, we revisit on-policy Reverse Kullback-Leibler (RKL) distillation and decompose its objective into a teacher-fitting term and a student-entropy term, without introducing an explicit FKL branch. We show theoretically that the token-level optimal student distribution corresponds to a tempered variant of the teacher distribution, where the adaptive weight controls the trade-off between mode-seeking and uncertainty preservation. Guided by this insight, we propose \textbf{Adaptive Entropy Distillation (AED)}, which uses the teacher's entropy to dynamically calibrate token-level imitation strength. Experiments on instruction-following and mathematical reasoning benchmarks demonstrate that AED achieves superior overall performance and generally improves teacher--student distributional and entropy alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。