arXiv:2503.18626cs.CV2025-03ECCV

用扩散模型生成精简数据集,兼顾多样性与图像质量。

Generative Dataset Distillation using Min-Max Diffusion Model

  • 以扩散模型为生成器,通过极小极大损失控制数据集多样性
  • 提出采样步数缩减策略,在保证质量前提下提升生成效率
  • 在ECCV2024数据集蒸馏挑战赛中获生成赛道第二名

本文研究基于生成模型的图像数据集蒸馏问题,利用扩散模型生成图像以构建替代数据集。生成器可在保持评估时间的前提下生成任意数量的图像。本工作采用主流扩散模型作为生成器,并引入极小极大损失,在训练过程中调控数据集的多样性和代表性。然而,扩散模型生成图像耗时较长,因其依赖迭代过程。我们观察到生成样本数量与图像质量之间存在关键权衡,该权衡由扩散步数控制,因此提出扩散步数缩减策略以实现最优性能。本文详细阐述了所提方法及其效果。我们的模型在首届ECCV2024数据集蒸馏挑战赛生成赛道中获得第二名,验证了其卓越性能。

原文摘要 · Abstract (English)

In this paper, we address the problem of generative dataset distillation that utilizes generative models to synthesize images. The generator may produce any number of images under a preserved evaluation time. In this work, we leverage the popular diffusion model as the generator to compute a surrogate dataset, boosted by a min-max loss to control the dataset's diversity and representativeness during training. However, the diffusion model is time-consuming when generating images, as it requires an iterative generation process. We observe a critical trade-off between the number of image samples and the image quality controlled by the diffusion steps and propose Diffusion Step Reduction to achieve optimal performance. This paper details our comprehensive method and its performance. Our model achieved $2^{nd}$ place in the generative track of \href{https://www.dd-challenge.com/#/}{The First Dataset Distillation Challenge of ECCV2024}, demonstrating its superior performance.

数据集蒸馏扩散模型生成式学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。