提出新评估框架,揭示数据蒸馏效果被误导的根源
DD-Ranking: Rethinking the Evaluation of Dataset Distillation
- 构建统一评估框架,剥离额外技术干扰
- 实验证明随机采样图像也能表现优异
- 适合关注评估公平性的数据压缩研究者
近年来,数据蒸馏为数据压缩提供了可靠方案,用合成小数据集训练的模型性能可媲美原始数据集。为提升合成数据质量,多种训练流程与优化目标被提出,推动该领域发展。近期解耦式数据蒸馏方法在后评估阶段引入软标签和更强数据增强,并将蒸馏规模扩展至大尺寸数据集(如 ImageNet-1K)。然而,这引发疑问:准确率是否仍是公平评估数据蒸馏方法的可靠指标?我们的实证发现,这些方法的性能提升往往源于附加技术而非图像本身的质量,甚至随机采样的图像也能取得更优结果。这种评估设置失衡严重阻碍了数据蒸馏的发展。为此,我们提出 DD-Ranking,一个统一的评估框架及新的通用评价指标,以揭示不同方法的真实性能提升。通过聚焦蒸馏数据的实际信息增益,DD-Ranking 提供了更全面、公平的未来研究评估标准。
原文摘要 · Abstract (English)
In recent years, dataset distillation has provided a reliable solution for data compression, where models trained on the resulting smaller synthetic datasets achieve performance comparable to those trained on the original datasets. To further improve the performance of synthetic datasets, various training pipelines and optimization objectives have been proposed, greatly advancing the field of dataset distillation. Recent decoupled dataset distillation methods introduce soft labels and stronger data augmentation during the post-evaluation phase and scale dataset distillation up to larger datasets (e.g., ImageNet-1K). However, this raises a question: Is accuracy still a reliable metric to fairly evaluate dataset distillation methods? Our empirical findings suggest that the performance improvements of these methods often stem from additional techniques rather than the inherent quality of the images themselves, with even randomly sampled images achieving superior results. Such misaligned evaluation settings severely hinder the development of DD. Therefore, we propose DD-Ranking, a unified evaluation framework, along with new general evaluation metrics to uncover the true performance improvements achieved by different methods. By refocusing on the actual information enhancement of distilled datasets, DD-Ranking provides a more comprehensive and fair evaluation standard for future research advancements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。