arXiv:2502.17709cs.CVcs.AI2025-02ICML被引 3

通过对比生成合成数据,提升大模型对罕见概念的识别能力

Contrastive Visual Data Augmentation

论文配图:Contrastive Visual Data Augmentation
图 1 · 摘自论文原文
  • 针对易混淆概念提取视觉与语言对比特征,生成针对性合成数据
  • 在NovelSpecies等数据集上准确率提升最高达12.3%
  • 适合需要增强罕见或新概念识别能力的研究者使用

大型多模态模型(LMMs)常因依赖预训练知识且难以捕捉细微视觉特征,而难以识别新概念。训练数据中的领域知识缺失也导致其容易混淆视觉相似、常见误标或低资源概念。为此,我们提出对比视觉数据增强(CoDA)策略:提取目标概念与易混淆概念之间的关键对比文本和视觉特征,并利用多模态生成模型生成针对性合成数据。通过自动过滤机制保证特征与图像质量,经人工标注验证。在INaturalist、SUN等低资源概念与多样场景识别数据集上验证了CoDA的有效性与高效性。此外,我们构建了NovelSpecies基准数据集,包含大模型从未见过的新发现动物物种。在三组数据集上,LLaVA-1.6进行单样本更新的结果显示,CoDA相较当前最优数据增强方法在准确率上分别提升12.3%(NovelSpecies)、5.1%(SUN)和6.0%(iNat)。

原文摘要 · Abstract (English)

Large multimodal models (LMMs) often struggle to recognize novel concepts, as they rely on pre-trained knowledge and have limited ability to capture subtle visual details. Domain-specific knowledge gaps in training also make them prone to confusing visually similar, commonly misrepresented, or low-resource concepts. To help LMMs better align nuanced visual features with language, improving their ability to recognize and reason about novel or rare concepts, we propose a Contrastive visual Data Augmentation (CoDA) strategy. CoDA extracts key contrastive textual and visual features of target concepts against the known concepts they are misrecognized as, and then uses multimodal generative models to produce targeted synthetic data. Automatic filtering of extracted features and augmented images is implemented to guarantee their quality, as verified by human annotators. We show the effectiveness and efficiency of CoDA on low-resource concept and diverse scene recognition datasets including INaturalist and SUN. We additionally collect NovelSpecies, a benchmark dataset consisting of newly discovered animal species that are guaranteed to be unseen by LMMs. LLaVA-1.6 1-shot updating results on these three datasets show CoDA significantly improves SOTA visual data augmentation strategies by 12.3% (NovelSpecies), 5.1% (SUN), and 6.0% (iNat) absolute gains in accuracy.

数据增强多模态罕见概念合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。