通过生成细微差异的负样本,提升模型对视觉语言细节的理解能力。
Enhancing Fine-Grained Vision-Language Pretraining with Negative Augmented Samples
- 用视觉词典构建跨模态语义桥梁,生成仅在令牌层面有差异的负样本。
- 在多个细粒度任务上显著优于基线模型,尤其在区分相似图像时表现更优。
- 适合需要精准理解图文细微差别的应用,如医学图像标注、商品细节检索。
现有视觉语言预训练(VLP)方法在多种任务上取得显著进展,验证了其捕捉粗粒度语义关联的有效性。然而,其在细粒度理解方面的能力仍受限。主流VLP模型常忽略不同模态特征表达的精细差异,依赖整体特征相似性进行跨模态交互,且直接对齐与融合多模态特征,侧重于粗粒度通用表征,难以捕捉任务所需的细微感知差异。针对此问题,本文提出负向增强样本(NAS),一种专为提升细粒度理解而设计的视觉语言预训练模型。NAS利用视觉词典(VD)作为视觉与语言域间的语义桥梁,并基于VD提出负向视觉增强(NVA)方法,生成仅在令牌级别与正样本存在差异的挑战性负样本。这迫使模型以更高精度识别正负样本之间的微小差别。大量实验证明,NAS各组件有效,显著增强了细粒度视觉语言理解能力。
原文摘要 · Abstract (English)
Existing Vision-Language Pretraining (VLP) methods have achieved remarkable improvements across a variety of vision-language tasks, confirming their effectiveness in capturing coarse-grained semantic correlations. However, their capability for fine-grained understanding, which is critical for many nuanced vision-language applications, remains limited. Prevailing VLP models often overlook the intricate distinctions in expressing different modal features and typically depend on the similarity of holistic features for cross-modal interactions. Moreover, these models directly align and integrate features from different modalities, focusing more on coarse-grained general representations, thus failing to capture the nuanced differences necessary for tasks demanding a more detailed perception. In response to these limitations, we introduce Negative Augmented Samples(NAS), a refined vision-language pretraining model that innovatively incorporates NAS to specifically address the challenge of fine-grained understanding. NAS utilizes a Visual Dictionary(VD) as a semantic bridge between visual and linguistic domains. Additionally, it employs a Negative Visual Augmentation(NVA) method based on the VD to generate challenging negative image samples. These samples deviate from positive samples exclusively at the token level, thereby necessitating that the model discerns the subtle disparities between positive and negative samples with greater precision. Comprehensive experiments validate the efficacy of NAS components and underscore its potential to enhance fine-grained vision-language comprehension.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。