arXiv:2605.06157cs.CLcs.AI2026-05被引 3

用难负样本标题提升视觉语言模型细粒度理解能力

HNC: Leveraging Hard Negative Captions towards Models with Fine-Grained Visual-Linguistic Comprehension Capabilities

论文配图:HNC: Leveraging Hard Negative Captions towards Models with Fine-Grained Visual-Linguistic Comprehension Capabilities
图 1 · 摘自论文原文
  • 构建自动构造的难负文本数据集,强化跨模态匹配
  • 在零样本匹配检测任务上显著提升模型性能
  • 适合需要强跨模态推理能力的研究者使用

图像-文本匹配(ITM)是视觉与语言(VL)领域学习通用表征的主流方法。然而,由于网络收集的图文对关联较弱,模型难以实现模态间细粒度语义融合。为此,本文提出硬负样本标题(HNC):一个自动生成的数据集,包含用于ITM训练的刻意误导性难负文本,以促进细粒度跨模态理解。同时,我们构建了一个手动设计的挑战性测试集,用于评估模型在不同组合复杂度下的细粒度跨模态不匹配识别能力。实验表明,基于HNC训练可显著提升模型在诊断性任务上的零样本匹配检测能力,并在噪声视觉输入下表现更鲁棒。此外,HNC模型能提供与现有方法相当或更优的微调初始化效果。

原文摘要 · Abstract (English)

Image-Text-Matching (ITM) is one of the defacto methods of learning generalized representations from a large corpus in Vision and Language (VL). However, due to the weak association between the web-collected image-text pairs, models fail to show a fine-grained understanding of the combined semantics of these modalities. To address this issue we propose Hard Negative Captions (HNC): an automatically created dataset containing foiled hard negative captions for ITM training towards achieving fine-grained cross-modal comprehension in VL. Additionally, we provide a challenging manually-created test set for benchmarking models on a fine-grained cross-modal mismatch task with varying levels of compositional complexity. Our results show the effectiveness of training on HNC by improving the models' zero-shot capabilities in detecting mismatches on diagnostic tasks and performing robustly under noisy visual input scenarios. Also, we demonstrate that HNC models yield a comparable or better initialization for fine-tuning

视觉语言跨模态负样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。