让生成数据更真实多样,同时精准匹配文本描述。
Salient Concept-Aware Generative Data Augmentation
- 用显著概念感知模型过滤无关视觉细节,提升图文对齐度。
- 在8个细粒度数据集上平均准确率提升0.73%(常规)和6.5%(长尾)。
- 适合需要高质量数据增强的图像分类与小样本学习场景。
近期基于图像和文本提示的生成式数据增强方法难以平衡保真度与多样性,因合成过程中表征常被环境等非关键视觉属性纠缠,导致与文本提示产生冲突。为此,我们提出一种个性化图像生成框架,利用显著概念感知图像嵌入模型,在合成过程中降低无关视觉细节的影响,从而保持图像与文本输入之间的直观一致性。通过生成更保留类别判别特征且具可控变化的图像,该框架有效提升了训练数据集的多样性,增强了下游模型的鲁棒性。在八个细粒度视觉数据集上,本方法优于当前最先进增强技术,常规设置下平均分类准确率提升0.73%,长尾设置下提升6.5%。
原文摘要 · Abstract (English)
Recent generative data augmentation methods conditioned on both image and text prompts struggle to balance between fidelity and diversity, as it is challenging to preserve essential image details while aligning with varied text prompts. This challenge arises because representations in the synthesis process often become entangled with non-essential input image attributes such as environmental contexts, creating conflicts with text prompts intended to modify these elements. To address this, we propose a personalized image generation framework that uses a salient concept-aware image embedding model to reduce the influence of irrelevant visual details during the synthesis process, thereby maintaining intuitive alignment between image and text inputs. By generating images that better preserve class-discriminative features with additional controlled variations, our framework effectively enhances the diversity of training datasets and thereby improves the robustness of downstream models. Our approach demonstrates superior performance across eight fine-grained vision datasets, outperforming state-of-the-art augmentation methods with averaged classification accuracy improvements by 0.73% and 6.5% under conventional and long-tail settings, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。