arXiv:2505.22946cs.CLcs.AI2025-05ACL被引 14

测试视觉语言模型对否定句的理解能力,发现普遍表现不佳。

NegVQA: Can Vision Language Models Understand Negation?

  • 用大模型生成带否定的视觉问答题,构建新基准
  • 20个主流模型在否定题上性能显著下降
  • 小模型和超大模型表现更好,中间规模最差

否定是能完全反转句子含义的语言现象。随着视觉语言模型(VLMs)持续发展并应用于高风险场景,评估其否定理解能力至关重要。为此,我们提出NegVQA,一个包含7,379个二选一问题的视觉问答基准,覆盖多样化的否定场景和图像-问题分布。通过利用大语言模型生成现有VQA数据集中问题的否定版本来构建该基准。我们在七个模型系列中评估了20个最先进的VLMs,发现这些模型在否定任务上表现显著不足,相比原问题响应有明显性能下降。此外,我们发现了U型缩放趋势:模型规模增大初期会降低负面对话性能,随后才逐步改善。该基准揭示了VLM在否定理解上的关键缺陷,并为未来模型发展提供重要洞见。

原文摘要 · Abstract (English)

Negation is a fundamental linguistic phenomenon that can entirely reverse the meaning of a sentence. As vision language models (VLMs) continue to advance and are deployed in high-stakes applications, assessing their ability to comprehend negation becomes essential. To address this, we introduce NegVQA, a visual question answering (VQA) benchmark consisting of 7,379 two-choice questions covering diverse negation scenarios and image-question distributions. We construct NegVQA by leveraging large language models to generate negated versions of questions from existing VQA datasets. Evaluating 20 state-of-the-art VLMs across seven model families, we find that these models struggle significantly with negation, exhibiting a substantial performance drop compared to their responses to the original questions. Furthermore, we uncover a U-shaped scaling trend, where increasing model size initially degrades performance on NegVQA before leading to improvements. Our benchmark reveals critical gaps in VLMs' negation understanding and offers insights into future VLM development. Project page available at https://yuhui-zh15.github.io/NegVQA/.

视觉语言模型否定理解基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。