解决扩散模型数据蒸馏中的目标与条件不一致问题,提升小样本数据集性能。
CaO$_2$: Rectifying Inconsistencies in Diffusion-Based Dataset Distillation
- 分两阶段优化:先选样本,再精调潜在表示以匹配条件
- 在ImageNet上平均准确率比基线高2.3%
- 适合需要高效压缩大图像数据集的研究者
近期将扩散模型引入数据蒸馏,展现出生成紧凑代理数据集的潜力,相比传统双层/单层优化方法更具效率和性能。然而现有方法忽视评估过程,在蒸馏中存在两大关键不一致:(1) 目标不一致——蒸馏过程偏离评估目标;(2) 条件不一致——生成图像与其对应条件不匹配。为解决这些问题,我们提出条件感知的目标引导采样框架CaO₂,通过两阶段扩散机制对齐蒸馏过程与评估目标。第一阶段采用概率感知的样本选择流程,第二阶段优化对应潜在表示以提升条件似然。CaO₂在ImageNet及其子集上达到当前最优性能,平均超越最佳基线2.3%准确率。
原文摘要 · Abstract (English)
The recent introduction of diffusion models in dataset distillation has shown promising potential in creating compact surrogate datasets for large, high-resolution target datasets, offering improved efficiency and performance over traditional bi-level/uni-level optimization methods. However, current diffusion-based dataset distillation approaches overlook the evaluation process and exhibit two critical inconsistencies in the distillation process: (1) Objective Inconsistency, where the distillation process diverges from the evaluation objective, and (2) Condition Inconsistency, leading to mismatches between generated images and their corresponding conditions. To resolve these issues, we introduce Condition-aware Optimization with Objective-guided Sampling (CaO$_2$), a two-stage diffusion-based framework that aligns the distillation process with the evaluation objective. The first stage employs a probability-informed sample selection pipeline, while the second stage refines the corresponding latent representations to improve conditional likelihood. CaO$_2$ achieves state-of-the-art performance on ImageNet and its subsets, surpassing the best-performing baselines by an average of 2.3% accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。