用扩散模型生成更多样数据,提升压缩数据集的泛化能力
Diversity-Driven Generative Dataset Distillation Based on Diffusion Model with Self-Adaptive Memory
- 通过自适应记忆机制评估数据分布相似性
- 生成数据多样性显著提升,验证准确率更高
- 适合需要高效训练的模型压缩场景
数据蒸馏可通过将大规模数据集压缩为小而具代表性的数据集,显著缩短深度神经网络的训练时间。尽管生成模型在该领域已取得进展,但其蒸馏数据的分布多样性不足,导致下游验证精度下降。本文提出一种基于扩散模型的多样性驱动生成式数据蒸馏方法,引入自适应记忆机制,评估蒸馏数据与真实数据之间的分布对齐程度,并据此引导扩散模型生成更具多样性的数据。大量实验表明,该方法在多数情况下优于现有最先进方法,证明其在数据蒸馏任务中的有效性。
原文摘要 · Abstract (English)
Dataset distillation enables the training of deep neural networks with comparable performance in significantly reduced time by compressing large datasets into small and representative ones. Although the introduction of generative models has made great achievements in this field, the distributions of their distilled datasets are not diverse enough to represent the original ones, leading to a decrease in downstream validation accuracy. In this paper, we present a diversity-driven generative dataset distillation method based on a diffusion model to solve this problem. We introduce self-adaptive memory to align the distribution between distilled and real datasets, assessing the representativeness. The degree of alignment leads the diffusion model to generate more diverse datasets during the distillation process. Extensive experiments show that our method outperforms existing state-of-the-art methods in most situations, proving its ability to tackle dataset distillation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。