arXiv:2601.01484cs.LG2026-01中稿 · ICLR被引 1

用贝叶斯教师提升学生模型训练稳定性与准确率

SGD-Based Knowledge Distillation with Bayesian Teachers: Theory and Guidelines

  • 从贝叶斯视角分析SGD训练下学生模型的收敛性
  • 贝叶斯教师使准确率提升最高达4.27%,噪声降低30%
  • 适合关注知识蒸馏理论与稳定训练的研究者

知识蒸馏(KD)是将大型教师网络的知识迁移至较小学生模型的核心方法,通常依赖软概率输出。尽管KD在众多应用中表现出色,其理论基础仍不完整。本文从贝叶斯角度分析基于随机梯度下降(SGD)的学生模型收敛行为,考察两种情形:(i) 教师提供精确的贝叶斯类别概率(BCPs);(ii) 使用噪声近似替代的BCPs。分析表明,使用真实BCPs可减少方差并消除收敛界中的邻域项,优于单热标签监督。我们进一步刻画了噪声水平对泛化与准确率的影响。基于此,建议使用能更好估计BCPs的贝叶斯深度学习模型作为教师。实验验证:相比确定性教师,来自贝叶斯教师的学生模型准确率最高提升4.27%,收敛过程噪声降低30%。

原文摘要 · Abstract (English)

Knowledge Distillation (KD) is a central paradigm for transferring knowledge from a large teacher network to a typically smaller student model, often by leveraging soft probabilistic outputs. While KD has shown strong empirical success in numerous applications, its theoretical underpinnings remain only partially understood. In this work, we adopt a Bayesian perspective on KD to rigorously analyze the convergence behavior of students trained with Stochastic Gradient Descent (SGD). We study two regimes: $(i)$ when the teacher provides the exact Bayes Class Probabilities (BCPs); and $(ii)$ supervision with noisy approximations of the BCPs. Our analysis shows that learning from BCPs yields variance reduction and removes neighborhood terms in the convergence bounds compared to one-hot supervision. We further characterize how the level of noise affects generalization and accuracy. Motivated by these insights, we advocate the use of Bayesian deep learning models, which typically provide improved estimates of the BCPs, as teachers in KD. Consistent with our analysis, we experimentally demonstrate that students distilled from Bayesian teachers not only achieve higher accuracies (up to +4.27%), but also exhibit more stable convergence (up to 30% less noise), compared to students distilled from deterministic teachers.

知识蒸馏贝叶斯学习模型训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。