改进扩散模型生成数据的分布覆盖,提升分类性能。
IMS3: Breaking Distributional Aggregation in Diffusion-Based Dataset Distillation
- 用反演对齐微调扩散过程,扩大样本分布范围。
- 无需训练采样高区分度子集,增强类间可分性。
- 适合需要高质量小规模数据集的研究者使用。
数据蒸馏旨在合成紧凑数据集以逼近大规模真实数据集的训练效果,缓解现代深度学习日益增长的计算需求。近期基于扩散模型的数据蒸馏方法利用扩散模型强大的生成能力,生成多样且结构一致的样本,展现出巨大潜力。然而,根本目标错位仍存:扩散模型优化的是生成似然而非判别效用,导致样本过度集中于高密度区域,边界样本覆盖不足,影响分类性能。为此,我们提出两种互补策略。反演匹配(IM)引入反演引导的微调过程,使去噪轨迹与反演路径对齐,拓宽分布覆盖并提升多样性。选择性子组采样(S^3)是一种无训练采样机制,通过选取兼具代表性和差异性的合成子集,提升类间可分性。大量实验表明,本方法显著提升了蒸馏数据集的判别质量和泛化能力,在基于扩散的方法中达到当前最优性能。
原文摘要 · Abstract (English)
Dataset Distillation aims to synthesize compact datasets that can approximate the training efficacy of large-scale real datasets, offering an efficient solution to the increasing computational demands of modern deep learning. Recently, diffusion-based dataset distillation methods have shown great promise by leveraging the strong generative capacity of diffusion models to produce diverse and structurally consistent samples. However, a fundamental goal misalignment persists: diffusion models are optimized for generative likelihood rather than discriminative utility, resulting in over-concentration in high-density regions and inadequate coverage of boundary samples crucial for classification. To address this issue, we propose two complementary strategies. Inversion-Matching (IM) introduces an inversion-guided fine-tuning process that aligns denoising trajectories with their inversion counterparts, broadening distributional coverage and enhancing diversity. Selective Subgroup Sampling(S^3) is a training-free sampling mechanism that improves inter-class separability by selecting synthetic subsets that are both representative and distinctive. Extensive experiments demonstrate that our approach significantly enhances the discriminative quality and generalization of distilled datasets, achieving state-of-the-art performance among diffusion-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。