arXiv:2410.01611cs.CVcs.AI2024-10

用额外辅助信息提升数据压缩效果,让小数据集更高效。

DRUPI: Dataset Reduction Using Privileged Information

  • 在数据压缩时合成特征或注意力标签等特权信息作为辅助监督
  • 实验显示中等区分度的特征标签能最好提升模型性能
  • 可无缝集成现有方法,在多个图像数据集上显著提效

数据凝缩(DC)旨在从大规模数据集中选取或提炼出小规模子集,同时保持目标任务的性能。现有方法主要关注以原始数据格式(输入与标签对)进行数据剪枝或生成。然而,在数据凝缩场景下,我们发现可以合成超出原始数据-标签对的信息作为额外学习目标,以促进模型训练。本文提出利用特权信息的数据凝缩方法(DCPI),在生成压缩数据集的同时,同步合成特征标签或注意力标签等特权信息,为模型训练提供辅助监督。研究发现,有效的特征标签需在区分度和多样性之间取得平衡,中等程度表现最优。在ImageNet-1K、CIFAR-10/100和Tiny ImageNet上的大量实验表明,DCPI可无缝融合现有数据凝缩方法,并带来显著性能提升。

原文摘要 · Abstract (English)

Dataset Condensation (DC) seeks to select or distill samples from large datasets into smaller subsets while preserving performance on target tasks. Existing methods primarily focus on pruning or synthesizing data in the same format as the original dataset, typically being the input data and corresponding labels. However, in DC settings, we find it is possible to synthesize more information beyond the data-label pair as an additional learning target to facilitate model training. In this paper, we introduce Dataset Condensation using Privileged Information (DCPI), which enriches DC by synthesizing privileged information alongside the reduced dataset. This privileged information can take the form of feature labels or attention labels, providing auxiliary supervision to improve model learning. Our findings reveal that effective feature labels must balance between being overly discriminative and excessively diverse, with a moderate level proves optimal for improving the reduced dataset's efficacy. Extensive experiments on ImageNet-1K, CIFAR-10/100 and Tiny ImageNet demonstrate that DCPI integrates seamlessly with existing dataset condensation methods, offering significant performance gains.

数据凝缩特权信息图像分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。