arXiv:2409.01679cs.CVcs.AI2024-09被引 2

让学生模型更聪明地学习老师的知识,提升分类准确率。

Adaptive Explicit Knowledge Transfer for Knowledge Distillation

  • 通过自适应机制同时学习显式和隐式知识
  • 在CIFAR-100和ImageNet上超越现有方法
  • 适合需要高效模型压缩的研究者

基于输出概率的知识蒸馏(KD)相比特征级蒸馏成本更低,但性能通常较差。最近研究发现,将教师模型对非目标类别的概率分布(即‘隐性(暗)知识’)有效传递给学生模型,可提升性能。通过梯度分析,我们首次揭示这实际上实现了对隐性知识学习的自适应控制。为此,我们提出一种新损失函数,使学生能以自适应方式同时学习显性知识(即教师对目标类别的置信度)和隐性知识。此外,我们提出将分类与蒸馏任务分离,以实现更有效的蒸馏和类别间关系建模。实验结果表明,所提方法——自适应显性知识迁移(AEKT)——在CIFAR-100和ImageNet数据集上均优于当前最优的蒸馏方法。

原文摘要 · Abstract (English)

Logit-based knowledge distillation (KD) for classification is cost-efficient compared to feature-based KD but often subject to inferior performance. Recently, it was shown that the performance of logit-based KD can be improved by effectively delivering the probability distribution for the non-target classes from the teacher model, which is known as `implicit (dark) knowledge', to the student model. Through gradient analysis, we first show that this actually has an effect of adaptively controlling the learning of implicit knowledge. Then, we propose a new loss that enables the student to learn explicit knowledge (i.e., the teacher's confidence about the target class) along with implicit knowledge in an adaptive manner. Furthermore, we propose to separate the classification and distillation tasks for effective distillation and inter-class relationship modeling. Experimental results demonstrate that the proposed method, called adaptive explicit knowledge transfer (AEKT) method, achieves improved performance compared to the state-of-the-art KD methods on the CIFAR-100 and ImageNet datasets.

知识蒸馏自适应学习模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。