用真实图像微调扩散模型,生成更逼真的增强数据。
Dataset Augmentation by Mixing Visual Concepts
- 用真实图像和新文本嵌入微调扩散模型,缩小生成与真实数据差距。
- 在多个分类任务上超越现有增强方法,提升模型性能。
- 适合需要高质量数据增强的计算机视觉研究者。
本文提出一种通过微调预训练扩散模型实现数据集增强的方法。使用文本条件生成图像常导致真实数据与生成图像间存在领域差异。我们提出一种微调方法,通过真实图像和新型文本嵌入来调整扩散模型。引入名为混合视觉概念(Mixing Visual Concepts, MVC)的独特流程,从图像标题生成新文本嵌入。MVC 能生成多样且贴近真实数据的图像,实现有效数据集增强。我们在多个基准分类任务上进行了全面的定性与定量评估,验证了生成图像在粗粒度和细粒度上的变化效果。所提方法在多项任务中优于当前最优的数据增强技术。
原文摘要 · Abstract (English)
This paper proposes a dataset augmentation method by fine-tuning pre-trained diffusion models. Generating images using a pre-trained diffusion model with textual conditioning often results in domain discrepancy between real data and generated images. We propose a fine-tuning approach where we adapt the diffusion model by conditioning it with real images and novel text embeddings. We introduce a unique procedure called Mixing Visual Concepts (MVC) where we create novel text embeddings from image captions. The MVC enables us to generate multiple images which are diverse and yet similar to the real data enabling us to perform effective dataset augmentation. We perform comprehensive qualitative and quantitative evaluations with the proposed dataset augmentation approach showcasing both coarse-grained and finegrained changes in generated images. Our approach outperforms state-of-the-art augmentation techniques on benchmark classification tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。