arXiv:2507.21335cs.CV2025-07被引 3

用对话原则测试视觉模型对提问异常的容忍度。

Analyzing the Sensitivity of Vision Language Models in Visual Question Answering

  • 基于格赖斯会话准则,给问题加修饰词测试模型响应
  • 三款主流模型在加修饰后表现均下降,说明敏感性强
  • 适合研究多模态模型鲁棒性与人类认知差异的学者

我们将视觉问答视为人与AI之间的(多模态)对话。本文从格赖斯会话准则的角度,探究视觉语言模型(VLMs)对对话规则违背的敏感性。尽管人类在面对违反格赖斯准则的情况时仍能理解,只需额外认知努力,但本研究旨在检验VLM是否具备类似能力。我们对人工构建的问题添加修饰词,并分析GPT-4o、Claude-3.5-Sonnet和Gemini-1.5-Flash在VQA v2.0数据集上的响应。结果表明,随着修饰词的引入,三款模型性能均持续下降,提示该方法是评估VLM局限性的有效路径。

原文摘要 · Abstract (English)

We can think of Visual Question Answering as a (multimodal) conversation between a human and an AI system. Here, we explore the sensitivity of Vision Language Models (VLMs) through the lens of cooperative principles of conversation proposed by Grice. Specifically, even when Grice's maxims of conversation are flouted, humans typically do not have much difficulty in understanding the conversation even though it requires more cognitive effort. Here, we study if VLMs are capable of handling violations to Grice's maxims in a manner that is similar to humans. Specifically, we add modifiers to human-crafted questions and analyze the response of VLMs to these modifiers. We use three state-of-the-art VLMs in our study, namely, GPT-4o, Claude-3.5-Sonnet and Gemini-1.5-Flash on questions from the VQA v2.0 dataset. Our initial results seem to indicate that the performance of VLMs consistently diminish with the addition of modifiers which indicates our approach as a promising direction to understand the limitations of VLMs.

视觉问答多模态模型认知偏差鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。