通过对抗性增强缓解未知数据分布偏移下的知识蒸馏失效问题
Improving Knowledge Distillation Under Unknown Covariate Shift Through Confidence-Guided Data Augmentation
- 基于扩散模型生成教师与学生分歧最大的难样本
- 在三个数据集上提升平均准确率和最差组表现
- 适合小样本场景下需鲁棒蒸馏的模型压缩任务
大规模基础模型在广泛数据集上训练后,在多个领域展现出强大的零样本能力。当数据和模型规模受限时,知识蒸馏已成为将基础模型知识迁移至小型学生网络的成熟方法。然而,蒸馏效果常因训练数据覆盖不足而受阻,导致训练与测试数据间存在协变量偏移,使学生模型利用虚假特征甚至产生捷径学习。本文提出一种新颖的基于扩散的数据增强策略,通过最大化教师与学生之间的预测分歧来生成挑战性样本,有效缓解协变量偏移问题。实验表明,相较于现有最先进的扩散基数据增强基线,本方法在CelebA-HQ、SpuCo Birds、BAR数据集上的样本平均准确率达到最佳或第二佳,并提升了最差组与平均组准确率,同时在存在协变量偏移的SpuCo ImageNet上降低了虚假特征得分。
原文摘要 · Abstract (English)
Large foundation models trained on extensive datasets demonstrate strong zero-shot capabilities in various domains. Knowledge distillation has become an established tool for transferring knowledge from foundation models to small student networks when data and model size are constrained. However, the efficacy of distillation is often hampered by limited training data coverage. This can result in a covariate shift between training and test data which in turn can lead the student to exploit spurious features or even shortcut learning. We address this problem by introducing a novel diffusion-based data augmentation strategy that generates images by maximizing the disagreement between the teacher and the student, effectively creating challenging samples that the student struggles with, thus mitigating the problem of covariate shift. Experiments demonstrate that, compared to state-of-the-art diffusion-based data augmentation baselines, our approach is best or second-best in sample mean accuracy and improves the worst group and mean group accuracy on CelebA-HQ, SpuCo Birds and BAR as well as the spurious score on Spurious ImageNet under covariate shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。