arXiv:2504.01757stat.MLcs.LG2025-04中稿 · as a conference pa…

统一特征知识蒸馏方法,提升模型压缩效果

KD$^{2}$M: A unifying framework for feature knowledge distillation

  • 提出统一框架KD²M,基于特征分布匹配实现知识迁移
  • 在多个视觉数据集上验证有效性,性能优于传统方法
  • 提供理论分析与指标对比,适合研究模型压缩者参考

知识蒸馏(KD)旨在将教师网络的知识迁移到学生网络。传统方法通常通过匹配输出预测来实现,而近期一些工作则提出通过匹配神经网络激活值的分布(即特征分布)来进行知识迁移,这一过程称为分布匹配。本文提出一个统一框架——通过分布匹配的知识蒸馏(KD²M),系统化地形式化该策略。主要贡献包括:一、总结了用于分布匹配的各类分布度量;二、在计算机视觉数据集上进行基准测试;三、推导出关于知识蒸馏的新理论结果。

原文摘要 · Abstract (English)

Knowledge Distillation (KD) seeks to transfer the knowledge of a teacher, towards a student neural net. This process is often done by matching the networks' predictions (i.e., their output), but, recently several works have proposed to match the distributions of neural nets' activations (i.e., their features), a process known as \emph{distribution matching}. In this paper, we propose an unifying framework, Knowledge Distillation through Distribution Matching (KD$^{2}$M), which formalizes this strategy. Our contributions are threefold. We i) provide an overview of distribution metrics used in distribution matching, ii) benchmark on computer vision datasets, and iii) derive new theoretical results for KD.

知识蒸馏特征匹配模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。