用生成模型合成细粒度图像,提升小样本分类效果。
SGIA: Enhancing Fine-Grained Visual Classification with Sequence Generative Image Augmentation
- 基于序列扩散模型生成多样姿态的逼真图像。
- 在CUB-200-2011上比之前最好方法高0.5%准确率。
- 适合做细粒度分类且数据少的研究者使用。
在细粒度视觉分类(FGVC)中,区分高度相似的子类别仍具挑战性,通常需要具有广泛变异性的数据集。获取和标注此类数据集困难且成本高昂,需专业知识识别细微差异。本文提出一种新方法——序列生成图像增强(SGIA),利用序列潜在扩散模型(SLDM)生成合成数据,并引入桥接迁移学习(BTL)机制,缩小真实与合成数据之间的域差距。该方法显著优于现有技术,生成更逼真的图像样本,提供超越传统刚性变换和风格变化的多样化姿态转换。在多个数据集、模型和训练策略下验证了增强数据的有效性,尤其在少样本学习场景中表现突出。在三个FGVC数据集上的基准测试中,本方法在真实度、多样性与表征质量方面均优于传统增强技术。本工作设定新基准,在CUB-200-2011数据集上分类准确率超过前一最佳模型0.5%,推动生成模型在FGVC数据增强中的应用。
原文摘要 · Abstract (English)
In Fine-Grained Visual Classification (FGVC), distinguishing highly similar subcategories remains a formidable challenge, often necessitating datasets with extensive variability. The acquisition and annotation of such FGVC datasets are notably difficult and costly, demanding specialized knowledge to identify subtle distinctions among closely related categories. Our study introduces a novel approach employing the Sequence Latent Diffusion Model (SLDM) for augmenting FGVC datasets, called Sequence Generative Image Augmentation (SGIA). Our method features a unique Bridging Transfer Learning (BTL) process, designed to minimize the domain gap between real and synthetically augmented data. This approach notably surpasses existing methods in generating more realistic image samples, providing a diverse range of pose transformations that extend beyond the traditional rigid transformations and style changes in generative augmentation. We demonstrate the effectiveness of our augmented dataset with substantial improvements in FGVC tasks on various datasets, models, and training strategies, especially in few-shot learning scenarios. Our method outperforms conventional image augmentation techniques in benchmark tests on three FGVC datasets, showcasing superior realism, variability, and representational quality. Our work sets a new benchmark and outperforms the previous state-of-the-art models in classification accuracy by 0.5% for the CUB-200-2011 dataset and advances the application of generative models in FGVC data augmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。