用双重增强提升师生模型的不变特征学习能力
Distilling Invariant Representations with Dual Augmentation
- 在师生模型上分别施加不同增强,促进不变特征学习
- 在CIFAR-100上实现与同架构知识蒸馏相当的性能
- 适合关注鲁棒特征提取和模型泛化能力的研究者
知识蒸馏(KD)广泛用于将大型精确模型(教师)的知识迁移到小型高效模型(学生)。近期方法通过引入因果解释来强化一致性,以提炼不变表示。本文在此基础上提出双重增强策略,分别在教师和学生模型上施加不同数据增强,推动学生捕获更鲁棒、可迁移的特征。该策略补充了不变因果蒸馏,确保学习到的表示在更广泛的数据变化和变换下保持稳定。在CIFAR-100上的大量实验验证了该方法的有效性,实现了与同架构知识蒸馏相当的竞争力。
原文摘要 · Abstract (English)
Knowledge distillation (KD) has been widely used to transfer knowledge from large, accurate models (teachers) to smaller, efficient ones (students). Recent methods have explored enforcing consistency by incorporating causal interpretations to distill invariant representations. In this work, we extend this line of research by introducing a dual augmentation strategy to promote invariant feature learning in both teacher and student models. Our approach leverages different augmentations applied to both models during distillation, pushing the student to capture robust, transferable features. This dual augmentation strategy complements invariant causal distillation by ensuring that the learned representations remain stable across a wider range of data variations and transformations. Extensive experiments on CIFAR-100 demonstrate the effectiveness of this approach, achieving competitive results in same-architecture KD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。