让学生模型复现教师的神经坍缩结构,提升知识蒸馏效果
Neural Collapse Inspired Knowledge Distillation
- 通过模仿教师网络的神经坍缩几何结构进行知识迁移
- 显著缩小师生差距,提升学生模型泛化能力
- 适合追求高精度蒸馏模型的研究者和工程师
现有知识蒸馏方法虽能使学生网络性能媲美教师,但师生间仍存在显著的知识差距,可能影响蒸馏效果。本文将神经坍缩(Neural Collapse, NC)结构引入知识蒸馏框架。NC通常出现在训练末期,使最后一层特征形成等角紧框架的正单纯形结构,有助于提升深度网络泛化性。我们提出假设:NC可缓解蒸馏中的知识鸿沟,从而增强学生表现。通过实证分析,发现将教师的NC结构迁移到学生能有效促进蒸馏过程。因此,不同于以往仅传递实例级输出或特征的方法,我们提出一种新范式——神经坍缩启发的知识蒸馏(NCKD),鼓励学生学习教师的NC结构。大量实验表明,NCKD简单高效,显著提升所有学生模型的泛化性能,并达到当前最优准确率。
原文摘要 · Abstract (English)
Existing knowledge distillation (KD) methods have demonstrated their ability in achieving student network performance on par with their teachers. However, the knowledge gap between the teacher and student remains significant and may hinder the effectiveness of the distillation process. In this work, we introduce the structure of Neural Collapse (NC) into the KD framework. NC typically occurs in the final phase of training, resulting in a graceful geometric structure where the last-layer features form a simplex equiangular tight frame. Such phenomenon has improved the generalization of deep network training. We hypothesize that NC can also alleviate the knowledge gap in distillation, thereby enhancing student performance. This paper begins with an empirical analysis to bridge the connection between knowledge distillation and neural collapse. Through this analysis, we establish that transferring the teacher's NC structure to the student benefits the distillation process. Therefore, instead of merely transferring instance-level logits or features, as done by existing distillation methods, we encourage students to learn the teacher's NC structure. Thereby, we propose a new distillation paradigm termed Neural Collapse-inspired Knowledge Distillation (NCKD). Comprehensive experiments demonstrate that NCKD is simple yet effective, improving the generalization of all distilled student models and achieving state-of-the-art accuracy performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。