arXiv:2502.04475cs.CVcs.AI2025-02

用真实图像增强条件,让生成图像更适合作为训练数据。

Augmented Conditioning Is Enough For Effective Training Image Generation

  • 用真实图像+文本提示作为生成条件,提升多样性。
  • 在五个长尾和少样本分类任务上表现优于现有方法。
  • 无需微调模型,适合快速构建有效合成数据集。

文本到图像扩散模型的图像生成能力已显著提升,能从描述性文本生成高度逼真的图像,从而增强利用合成图像训练计算机视觉模型的可行性。为有效用于下游训练,生成图像需兼具高真实感与目标数据分布内的充分多样性。然而,当前最先进的条件图像生成模型主要针对创意应用优化,侧重图像真实性和提示遵循性,而忽视条件多样性。本文研究如何在不微调生成模型的前提下,通过增强条件来提升生成图像的多样性,以提高其对下游图像分类模型训练的有效性。我们发现,将生成过程同时基于增强的真实图像和文本提示进行条件化,可产生适合作为合成训练数据的图像。真实图像提供领域内上下文,使生成图像贴近真实数据分布;数据增强则引入视觉多样性,显著提升下游分类器性能。我们在五个经典长尾与少样本图像分类基准上验证该方法,结果表明,在长尾基准上持续优于现有最佳方法,并在其余四个基准的极端少样本场景中取得显著提升。这为有效利用合成数据进行下游训练迈出重要一步。

原文摘要 · Abstract (English)

Image generation abilities of text-to-image diffusion models have significantly advanced, yielding highly photo-realistic images from descriptive text and increasing the viability of leveraging synthetic images to train computer vision models. To serve as effective training data, generated images must be highly realistic while also sufficiently diverse within the support of the target data distribution. Yet, state-of-the-art conditional image generation models have been primarily optimized for creative applications, prioritizing image realism and prompt adherence over conditional diversity. In this paper, we investigate how to improve the diversity of generated images with the goal of increasing their effectiveness to train downstream image classification models, without fine-tuning the image generation model. We find that conditioning the generation process on an augmented real image and text prompt produces generations that serve as effective synthetic datasets for downstream training. Conditioning on real training images contextualizes the generation process to produce images that are in-domain with the real image distribution, while data augmentations introduce visual diversity that improves the performance of the downstream classifier. We validate augmentation-conditioning on a total of five established long-tail and few-shot image classification benchmarks and show that leveraging augmentations to condition the generation process results in consistent improvements over the state-of-the-art on the long-tailed benchmark and remarkable gains in extreme few-shot regimes of the remaining four benchmarks. These results constitute an important step towards effectively leveraging synthetic data for downstream training.

图像生成合成数据少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。