通过增强标签提升数据蒸馏效果,准确率平均提升14.9%
Label-Augmented Dataset Distillation
- 为合成图像添加密集标签,捕捉更丰富语义
- 仅增加2.5%存储空间,准确率平均提升14.9%
- 兼容多种算法,提升跨架构鲁棒性
传统数据蒸馏主要关注图像表征,常忽略标签的重要作用。本文提出标签增强型数据蒸馏(LADD),通过为每个合成图像子采样生成额外密集标签,以捕获丰富语义。该方法在ImageNet子集上仅增加2.5%存储开销,带来显著性能提升,提供强学习信号。标签生成策略可与现有数据蒸馏方法结合,显著提升训练效率与性能。实验表明,LADD在计算开销和准确率方面均优于现有方法;使用三种高性能蒸馏算法时,平均准确率提升达14.9%。该方法在不同数据集、蒸馏超参数及算法下均有效,并增强了蒸馏数据集的跨架构鲁棒性,对实际应用具有重要意义。
原文摘要 · Abstract (English)
Traditional dataset distillation primarily focuses on image representation while often overlooking the important role of labels. In this study, we introduce Label-Augmented Dataset Distillation (LADD), a new dataset distillation framework enhancing dataset distillation with label augmentations. LADD sub-samples each synthetic image, generating additional dense labels to capture rich semantics. These dense labels require only a 2.5% increase in storage (ImageNet subsets) with significant performance benefits, providing strong learning signals. Our label generation strategy can complement existing dataset distillation methods for significantly enhancing their training efficiency and performance. Experimental results demonstrate that LADD outperforms existing methods in terms of computational overhead and accuracy. With three high-performance dataset distillation algorithms, LADD achieves remarkable gains by an average of 14.9% in accuracy. Furthermore, the effectiveness of our method is proven across various datasets, distillation hyperparameters, and algorithms. Finally, our method improves the cross-architecture robustness of the distilled dataset, which is important in the application scenario.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。