arXiv:2409.01389cs.CL2024-09

提出新数据集,揭示视觉语言模型在理解依赖语境的动词时存在短板。

CV-Probes: Studying the interplay of lexical and world knowledge in visually grounded verb understanding

  • 构建包含情境依赖动词的图文配对数据集
  • 模型对需社会知识的动词理解准确率显著下降
  • 适合研究视觉语言模型认知机制的学者

视觉语言(VL)Transformer 模型如何理解动词短语?它们是否整合了上下文与世界知识?我们引入了 CV-Probes 数据集,包含需要社会知识和视觉上下文才能解释的动词短语(如“乞求”),以及仅凭图像信息即可理解的动词短语(如“坐”)。结果表明,当动词高度依赖上下文时,模型表现明显下降。通过可解释性分析发现,模型对标题中的动词标记关注度不足。这些发现提示需改进 VL 模型的训练与评估方法。代码与数据集将公开于 https://github.com/ivana-13/CV-Probes。

原文摘要 · Abstract (English)

How do vision-language (VL) transformer models ground verb phrases and do they integrate contextual and world knowledge in this process? We introduce the CV-Probes dataset, containing image-caption pairs involving verb phrases that require both social knowledge and visual context to interpret (e.g., "beg"), as well as pairs involving verb phrases that can be grounded based on information directly available in the image (e.g., "sit"). We show that VL models struggle to ground VPs that are strongly context-dependent. Further analysis using explainable AI techniques shows that such models may not pay sufficient attention to the verb token in the captions. Our results suggest a need for improved methodologies in VL model training and evaluation. The code and dataset will be available https://github.com/ivana-13/CV-Probes.

视觉语言动词理解可解释性数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。