用确定性微分方程快速生成高质量小数据集,提升训练效率。
Path-Guided Flow Matching for Dataset Distillation
- 基于冻结VAE的隐空间流匹配,实现快速确定性合成。
- 仅需少量步骤即可达到78%模式覆盖率,效率提升7.6倍。
- 适合对数据压缩效率要求高的模型训练场景。
数据集蒸馏将大规模数据集压缩为紧凑的合成数据集,以在训练模型时保持相近性能。尽管基于扩散的方法已有进展,但这类方法通常依赖启发式引导或原型分配,存在采样耗时和轨迹不稳定问题,尤其在强控制或低每类样本数(IPC)条件下影响下游泛化能力。我们提出首个基于流匹配的生成式蒸馏框架——路径引导流匹配(PGFM),通过在少数步骤内求解常微分方程(ODE),实现快速确定性合成。PGFM在冻结变分自编码器(VAE)的隐空间中进行流匹配,学习从高斯噪声到数据分布的类别条件迁移。特别地,我们设计了一种连续路径-原型引导算法,确保轨迹稳定到达指定原型,同时保持多样性和效率。在高分辨率基准上的大量实验表明,PGFM在更少采样步骤下表现优于或媲美现有扩散基蒸馏方法,且效率显著提升,例如相比扩散基方法效率提高7.6倍,同时达到78%的模式覆盖率。
原文摘要 · Abstract (English)
Dataset distillation compresses large datasets into compact synthetic sets with comparable performance in training models. Despite recent progress on diffusion-based distillation, this type of method typically depends on heuristic guidance or prototype assignment, which comes with time-consuming sampling and trajectory instability and thus hurts downstream generalization especially under strong control or low IPC. We propose \emph{Path-Guided Flow Matching (PGFM)}, the first flow matching-based framework for generative distillation, which enables fast deterministic synthesis by solving an ODE in a few steps. PGFM conducts flow matching in the latent space of a frozen VAE to learn class-conditional transport from Gaussian noise to data distribution. Particularly, we develop a continuous path-to-prototype guidance algorithm for ODE-consistent path control, which allows trajectories to reliably land on assigned prototypes while preserving diversity and efficiency. Extensive experiments across high-resolution benchmarks demonstrate that PGFM matches or surpasses prior diffusion-based distillation approaches with fewer steps of sampling while delivering competitive performance with remarkably improved efficiency, e.g., 7.6$\times$ more efficient than the diffusion-based counterparts with 78\% mode coverage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。