多教师知识蒸馏新方法,提升模型准确率与不确定性评估
Multi-Teacher Knowledge Distillation via Teacher-Informed Mixture Priors

- 基于贝叶斯框架融合多个教师模型的先验知识
- 通过熵权重自适应调整教师贡献,提升泛化性能
- 适合需要可靠预测与不确定性的实际应用
知识蒸馏是模型压缩的有效方法,可高效部署复杂深度学习模型(如大语言模型)。然而其内在统计机制仍不明确,且在需要多样教师专长的真实场景中常忽视不确定性评估。为此,我们提出多教师贝叶斯知识蒸馏(MT-BKD),在贝叶斯框架下让学生模型从多个教师中学习。该方法利用贝叶斯推断捕捉蒸馏过程中的固有不确定性,引入教师信息先验,融合教师模型与任务特定训练数据的外部知识,提升泛化性、鲁棒性与可扩展性。此外,基于熵的权重机制自适应调整各教师影响,使学生有效整合多源专家知识。实验在合成数据和真实任务(包括蛋白质亚细胞定位预测与图像分类)上验证了该方法,结果表明性能提升且不确定性量化更可靠,充分展现MT-BKD的优势。
原文摘要 · Abstract (English)
Knowledge distillation is a powerful method for model compression, enabling the efficient deployment of complex deep learning models (teachers), including large language models. However, its underlying statistical mechanisms remain unclear, and uncertainty evaluation is often overlooked, especially in real-world scenarios requiring diverse teacher expertise. To address these challenges, we introduce \textit{Multi-Teacher Bayesian Knowledge Distillation} (MT-BKD), where a distilled student model learns from multiple teachers within the Bayesian framework. Our approach leverages Bayesian inference to capture inherent uncertainty in the distillation process. We introduce a teacher-informed prior, integrating external knowledge from teacher models and task-specific training data, offering better generalization, robustness, and scalability. Additionally, an entropy-based weighting mechanism adaptively adjusts each teacher's influence, allowing the student to combine multiple sources of expertise effectively. MT-BKD enhances the interpretability of the student model's learning process, improves predictive accuracy, and provides uncertainty quantification. We validate MT-BKD on both synthetic and real-world tasks, including protein subcellular location prediction and image classification. Our experiments show improved performance and robust uncertainty quantification, highlighting the strengths of our MT-BKD framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。