从信息论角度揭示知识蒸馏如何提升模型泛化能力
On the Generalization of Knowledge Distillation: An Information-Theoretic View
- 将师生训练建模为耦合随机过程,定义蒸馏分歧度量
- 推导出学生模型的泛化误差上下界,与蒸馏分歧相关
- 发现教师模型局部平坦性可显著收紧泛化上界
知识蒸馏广泛用于提升模型泛化性能,但其理论机制仍不清晰。本文将教师与学生模型的训练视为耦合的随机过程,引入蒸馏分歧(distillation divergence),即两者随机核之间的KL散度。在此框架下,我们推导了学生模型相对于教师泛化差距的两个泛化界:在子高斯假设下通过算法稳定性得到上界;在中心条件假设下得到更紧的下界,且对蒸馏分歧依赖更强。进一步提出一个考虑损失尖锐性的界,明确给出了紧致性成立的条件,并证明教师模型的局部平坦性可严格收紧该界。在线性高斯案例中,蒸馏分歧可分解为偏差、方差和秩瓶颈三部分成本,为蒸馏设计提供可解释的指导。
原文摘要 · Abstract (English)
Knowledge distillation is widely used to improve generalization in practice, yet its theoretical understanding remains elusive. In the standard distillation setting, a teacher model provides soft predictions to guide the training of a student model. We model teacher and student training as coupled stochastic processes and introduce a distillation divergence, defined as the Kullback-Leibler divergence between these two stochastic kernels. Within this framework, we derive two generalization bounds for the student model relative to the teacher's generalization gap: an upper bound under a sub-Gaussian assumption via algorithmic stability, and a lower bound under a central condition with sharper dependence on the distillation divergence. We further develop a loss-sharpness-aware bound with an explicit tightness regime, showing that the teacher's local flatness can strictly tighten the bound. Additionally, in a linear Gaussian case study, the distillation divergence admits an interpretable decomposition into bias, variance, and rank-bottleneck costs, yielding practical guidance for distillation design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。