提出不对称决策机制,让学生模型更好学习教师特征。
Asymmetric Decision-Making in Online Knowledge Distillation:Unifying Consensus and Divergence
- 学生侧重学共性特征,教师突出差异特征
- 在多个任务上超越现有方法,提升显著
- 适合做知识蒸馏的开发者快速部署
在线知识蒸馏(OKD)将知识迁移过程简化为单阶段训练,无需预训练教师模型。本文分析师生模型的中间空间特征,发现:(1)师生间相似特征主要集中在前景物体上;(2)教师对前景的关注程度高于学生。基于此,提出不对称决策机制(ADM),让学生模型聚焦高共识特征以增强一致性,同时让教师模型强调低相似区域以保持多样性,从而引导学生模型更接近教师的特征学习能力。大量实验表明,ADM在多种在线知识蒸馏设置下均优于现有方法,并在离线蒸馏、语义分割和扩散蒸馏任务中也表现优异。
原文摘要 · Abstract (English)
Online Knowledge Distillation (OKD) methods streamline the distillation training process into a single stage, eliminating the need for knowledge transfer from a pretrained teacher network to a more compact student network. This paper presents an innovative approach to leverage intermediate spatial representations. Our analysis of the intermediate features from both teacher and student models reveals two pivotal insights: (1) the similar features between students and teachers are predominantly focused on foreground objects. (2) teacher models emphasize foreground objects more than students. Building on these findings, we propose Asymmetric Decision-Making (ADM) to enhance feature consensus learning for student models while continuously promoting feature diversity in teacher models. Specifically, Consensus Learning for student models prioritizes spatial features with high consensus relative to teacher models. Conversely, Divergence Learning for teacher models highlights spatial features with lower similarity compared to student models, indicating superior performance by teacher models in these regions. Consequently, ADM facilitates the student models to catch up with the feature learning process of the teacher models. Extensive experiments demonstrate that ADM consistently surpasses existing OKD methods across various online knowledge distillation settings and also achieves superior results when applied to offline knowledge distillation, semantic segmentation and diffusion distillation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。