综述数据集蒸馏新进展,聚焦大规模高效压缩与泛化能力提升。
The Evolution of Dataset Distillation: Toward Scalable and Generalizable Solutions
- 按轨迹、梯度、分布匹配等思路分类近年方法
- 提出SRe2L框架与软标签策略,显著提升压缩效率与模型精度
- 适合关注高效训练、跨领域应用的研究者
数据集蒸馏通过将大规模数据集压缩为紧凑的合成表示,成为高效训练现代深度学习模型的关键技术。现有综述多覆盖2023年前进展,本文系统梳理近期突破,重点涵盖对ImageNet-1K和ImageNet-21K等大规模数据集的可扩展性。方法分为轨迹匹配、梯度匹配、分布匹配、可扩展生成方法及解耦优化机制。突出创新包括:SRe2L框架实现高效精准压缩、软标签策略显著提升模型准确率、无损蒸馏技术在保持性能前提下最大化压缩。此外,探讨了对抗攻击、后门攻击下的鲁棒性,以及非独立同分布(non-IID)数据处理挑战。拓展至视频、音频、多模态学习、医学影像与科学计算等新兴应用,展现其广泛适用性。通过全面性能对比与研究方向建议,为研究人员提供实用指导,推动高效且通用的数据集蒸馏发展。
原文摘要 · Abstract (English)
Dataset distillation, which condenses large-scale datasets into compact synthetic representations, has emerged as a critical solution for training modern deep learning models efficiently. While prior surveys focus on developments before 2023, this work comprehensively reviews recent advances, emphasizing scalability to large-scale datasets such as ImageNet-1K and ImageNet-21K. We categorize progress into a few key methodologies: trajectory matching, gradient matching, distribution matching, scalable generative approaches, and decoupling optimization mechanisms. As a comprehensive examination of recent dataset distillation advances, this survey highlights breakthrough innovations: the SRe2L framework for efficient and effective condensation, soft label strategies that significantly enhance model accuracy, and lossless distillation techniques that maximize compression while maintaining performance. Beyond these methodological advancements, we address critical challenges, including robustness against adversarial and backdoor attacks, effective handling of non-IID data distributions. Additionally, we explore emerging applications in video and audio processing, multi-modal learning, medical imaging, and scientific computing, highlighting its domain versatility. By offering extensive performance comparisons and actionable research directions, this survey equips researchers and practitioners with practical insights to advance efficient and generalizable dataset distillation, paving the way for future innovations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。