教师模型的校准程度直接影响学生模型性能,改进校准可提升知识蒸馏效果。
The Role of Teacher Calibration in Knowledge Distillation
- 通过降低教师模型的校准误差来提升知识蒸馏效果
- 在分类与检测任务中均实现性能提升,优于现有方法
- 可无缝集成到主流方法中,适用性强
知识蒸馏(KD)是深度学习中一种有效的模型压缩技术,能够将大型教师模型的知识迁移到小型学生模型。尽管KD已取得显著成功,但其影响学生模型性能的关键因素仍不明确。本文揭示了教师模型的校准误差与学生模型准确率之间存在强相关性,表明教师校准是有效知识蒸馏的重要因素。进一步地,我们证明仅通过采用能降低教师校准误差的校准方法,即可提升KD性能。该算法具有通用性,在从分类到检测等多种任务上均有效,且可轻松融入现有最先进方法,持续取得更优结果。
原文摘要 · Abstract (English)
Knowledge Distillation (KD) has emerged as an effective model compression technique in deep learning, enabling the transfer of knowledge from a large teacher model to a compact student model. While KD has demonstrated significant success, it is not yet fully understood which factors contribute to improving the student's performance. In this paper, we reveal a strong correlation between the teacher's calibration error and the student's accuracy. Therefore, we claim that the calibration of the teacher model is an important factor for effective KD. Furthermore, we demonstrate that the performance of KD can be improved by simply employing a calibration method that reduces the teacher's calibration error. Our algorithm is versatile, demonstrating effectiveness across various tasks from classification to detection. Moreover, it can be easily integrated with existing state-of-the-art methods, consistently achieving superior performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。