用扩散模型实现无损数据浓缩,大幅减小数据集体积且不损失性能
Beyond Dataset Distillation: Lossless Dataset Concentration via Diffusion-Assisted Distribution Alignment
- 基于扩散模型的噪声优化方法合成高代表性小数据集
- 在低数据量下达到最优性能,高数据量时压缩近一半仍零性能损失
- 支持有数据和无数据场景,适用于大规模视觉系统训练
大规模视觉识别系统的开发受限于数据集的高昂成本与获取难度。数据浓缩通过生成紧凑的代理数据集,提升训练、存储、迁移与隐私保护效率。现有基于扩散模型的数据浓缩方法存在理论依据不足、高数据量扩展效率低、无法在无数据场景使用等问题。本文建立理论框架,证明数据浓缩等价于分布匹配,并揭示该范式固有的效率极限。提出数据浓缩(DsCo)框架,采用扩散驱动的噪声优化(NOpt)方法生成少量但具代表性的样本,并可选通过‘掺杂’(Doping)将原始数据中精选样本与合成样本混合,突破数据浓缩的效率瓶颈。DsCo在有数据与无数据场景均适用,在低数据量下达当前最优性能,高数据量下可将数据集规模几乎减半而性能无损。
原文摘要 · Abstract (English)
The high cost and accessibility problem associated with large datasets hinder the development of large-scale visual recognition systems. Dataset Distillation addresses these problems by synthesizing compact surrogate datasets for efficient training, storage, transfer, and privacy preservation. The existing state-of-the-art diffusion-based dataset distillation methods face three issues: lack of theoretical justification, poor efficiency in scaling to high data volumes, and failure in data-free scenarios. To address these issues, we establish a theoretical framework that justifies the use of diffusion models by proving the equivalence between dataset distillation and distribution matching, and reveals an inherent efficiency limit in the dataset distillation paradigm. We then propose a Dataset Concentration (DsCo) framework that uses a diffusion-based Noise-Optimization (NOpt) method to synthesize a small yet representative set of samples, and optionally augments the synthetic data via "Doping", which mixes selected samples from the original dataset with the synthetic samples to overcome the efficiency limit of dataset distillation. DsCo is applicable in both data-accessible and data-free scenarios, achieving SOTA performances for low data volumes, and it extends well to high data volumes, where it nearly reduces the dataset size by half with no performance degradation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。