arXiv:2508.10631cs.CV2025-08NeurIPS被引 6

用少量真实图像提升生成图像多样性,兼顾质量与实用性。

Increasing the Utility of Synthetic Images through Chamfer Guidance

  • 通过少量真实样本引导生成,自动评估并优化合成图像分布
  • 仅需2张真实图即达96.4%精度、86.4%覆盖度,32张时提升至97.5%和92.7%
  • 无需无条件模型,采样效率提高31%,适合数据增强与小样本场景

条件图像生成模型有望生成无限合成训练数据。然而,近期生成质量的提升以牺牲多样性为代价,限制了其作为合成数据源的实用性。尽管已有基于引导的方法试图在质量或多样性上改进,但其(隐式或显式)效用函数常忽略合成与真实数据间的分布偏移。本文提出Chamfer Guidance:一种无需训练的引导方法,仅需少量真实样本即可刻画合成数据的质量与多样性。实验表明,该方法在ImageNet-1k及标准地理多样性基准上,可在保持或提升生成质量的同时显著增强多样性。使用2张真实图像即达96.4%精度与86.4%分布覆盖度,增至32张时分别提升至97.5%与92.7%。将合成数据用于下游图像分类任务,可实现分布内性能提升最高15%,分布外提升最高16%。此外,本方法无需使用无条件模型,采样阶段减少31%计算量。

原文摘要 · Abstract (English)

Conditional image generative models hold considerable promise to produce infinite amounts of synthetic training data. Yet, recent progress in generation quality has come at the expense of generation diversity, limiting the utility of these models as a source of synthetic training data. Although guidance-based approaches have been introduced to improve the utility of generated data by focusing on quality or diversity, the (implicit or explicit) utility functions oftentimes disregard the potential distribution shift between synthetic and real data. In this work, we introduce Chamfer Guidance: a training-free guidance approach which leverages a handful of real exemplar images to characterize the quality and diversity of synthetic data. We show that by leveraging the proposed Chamfer Guidance, we can boost the diversity of the generations w.r.t. a dataset of real images while maintaining or improving the generation quality on ImageNet-1k and standard geo-diversity benchmarks. Our approach achieves state-of-the-art few-shot performance with as little as 2 exemplar real images, obtaining 96.4% in terms of precision, and 86.4% in terms of distributional coverage, which increase to 97.5% and 92.7%, respectively, when using 32 real images. We showcase the benefits of the Chamfer Guidance generation by training downstream image classifiers on synthetic data, achieving accuracy boost of up to 15% for in-distribution over the baselines, and up to 16% in out-of-distribution. Furthermore, our approach does not require using the unconditional model, and thus obtains a 31% reduction in FLOPs w.r.t. classifier-free-guidance-based approaches at sampling time.

图像生成数据增强引导生成多样性提升

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。