arXiv:2504.21478cs.CVcs.NE2025-04被引 2

提出新方法提升无数据知识蒸馏的模型泛化能力,让学生模型学得更通用。

CAE-DFKD: Bridging the Transferability Gap in Data-Free Knowledge Distillation

  • 在嵌入层设计类别感知生成器,改进传统图像级蒸馏思路
  • 在多个下游任务中表现优异,迁移性能显著优于现有方法
  • 适合需要高效、通用模型压缩的工业场景

无数据知识蒸馏(DFKD)可在不访问真实训练数据的情况下,将预训练教师网络的知识迁移到目标学生模型。现有方法主要关注特定数据集上的图像识别性能,常忽视所学表征的可迁移性。本文提出类别感知嵌入无数据知识蒸馏(CAE-DFKD),从嵌入层出发解决以往依赖图像级方法在直接应用于DFKD时泛化能力不足的问题。实验表明: extit{ extbf{i.)}} 通过改变生成器训练范式,获得显著效率优势; extit{ extbf{ii.)}} 在图像识别任务上达到与当前最优DFKD方法相当的性能; extit{ extbf{iii.)}} 在下游任务中展现出卓越的可迁移性,验证了其学习表征的有效性。

原文摘要 · Abstract (English)

Data-Free Knowledge Distillation (DFKD) enables the knowledge transfer from the given pre-trained teacher network to the target student model without access to the real training data. Existing DFKD methods focus primarily on improving image recognition performance on associated datasets, often neglecting the crucial aspect of the transferability of learned representations. In this paper, we propose Category-Aware Embedding Data-Free Knowledge Distillation (CAE-DFKD), which addresses at the embedding level the limitations of previous rely on image-level methods to improve model generalization but fail when directly applied to DFKD. The superiority and flexibility of CAE-DFKD are extensively evaluated, including: \textit{\textbf{i.)}} Significant efficiency advantages resulting from altering the generator training paradigm; \textit{\textbf{ii.)}} Competitive performance with existing DFKD state-of-the-art methods on image recognition tasks; \textit{\textbf{iii.)}} Remarkable transferability of data-free learned representations demonstrated in downstream tasks.

知识蒸馏无数据训练模型压缩迁移能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。