arXiv:2411.10693cs.CV2024-11被引 1

提出多视角对比logit蒸馏,提升模型压缩效率与精度

Multi-perspective Contrastive Logit Distillation

  • 从语义角度重新设计logit计算方式,突破传统KL散度局限
  • 在多个图像数据集上达到当前最优性能,超越复杂特征蒸馏方法
  • 特别适合Vision Transformer的高效蒸馏,训练速度更快

以往知识蒸馏研究中,logit蒸馏的重要性常被忽视。为重振logit蒸馏,本文基于logits的语义特性重新审视其计算方式,探索更高效的利用方法。Logits通常包含大量高层语义信息,但传统使用KL散度直接计算的方式未考虑其语义属性,也未能充分挖掘logits潜力。为此,我们提出一种新颖且高效的logit蒸馏方法——多视角对比logit蒸馏(MCLD),显著提升logit蒸馏的性能与效率。相比现有logit蒸馏方法及复杂特征蒸馏方法,MCLD在CIFAR-100、ImageNet、Tiny-ImageNet和STL-10等多个数据集上的图像分类与迁移学习任务中均达到顶尖水平。同时,MCLD展现出更优的训练效率,并在蒸馏Vision Transformers时表现突出,凸显其显著优势。本研究揭示了logits在知识蒸馏中的巨大潜力,为未来研究提供重要启示。

原文摘要 · Abstract (English)

In previous studies on knowledge distillation, the significance of logit distillation has frequently been overlooked. To revitalize logit distillation, we present a novel perspective by reconsidering its computation based on the semantic properties of logits and exploring how to utilize it more efficiently. Logits often contain a substantial amount of high-level semantic information; however, the conventional approach of employing logits to compute Kullback-Leibler (KL) divergence does not account for their semantic properties. Furthermore, this direct KL divergence computation fails to fully exploit the potential of logits. To address these challenges, we introduce a novel and efficient logit distillation method, Multi-perspective Contrastive Logit Distillation (MCLD), which substantially improves the performance and efficacy of logit distillation. In comparison to existing logit distillation methods and complex feature distillation methods, MCLD attains state-of-the-art performance in image classification, and transfer learning tasks across multiple datasets, including CIFAR-100, ImageNet, Tiny-ImageNet, and STL-10. Additionally, MCLD exhibits superior training efficiency and outstanding performance with distilling on Vision Transformers, further emphasizing its notable advantages. This study unveils the vast potential of logits in knowledge distillation and seeks to offer valuable insights for future research.

知识蒸馏视觉模型Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。