arXiv:2601.19202cs.CL2026-01Conference of the …被引 6

测试视觉语言模型在文本误导下的表现,发现它们易被错误文字欺骗。

Do Images Speak Louder than Words? Investigating the Effect of Textual Misinformation in VLMs

  • 构建冲突文本数据集,让文字与图像信息矛盾
  • 11个主流模型平均性能下降超48.2%
  • 揭示模型依赖文本、忽视图像的严重缺陷

视觉语言模型(VLMs)在视觉问答(VQA)任务中表现出强大的多模态推理能力,但其对文本误导的鲁棒性尚未充分研究。现有研究主要关注纯文本领域的误导,却未明确VLM如何权衡不同模态间的矛盾信息。为此,本文提出CONTEXT-VQA(即冲突文本)数据集,包含带有系统生成的误导性提示的图像-问题对,这些提示故意与视觉证据冲突。设计并执行了全面的评估框架,对11个最先进的VLMs进行基准测试。实验表明,这些模型极易受误导性文本影响,常忽视清晰的视觉证据而选择矛盾文本,仅一轮说服性对话后平均性能下降超过48.2%。研究揭示了当前VLMs的关键缺陷,强调亟需提升对抗文本操纵的鲁棒性。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have shown strong multimodal reasoning capabilities on Visual-Question-Answering (VQA) benchmarks. However, their robustness against textual misinformation remains under-explored. While existing research has studied the effect of misinformation in text-only domains, it is not clear how VLMs arbitrate between contradictory information from different modalities. To bridge the gap, we first propose the CONTEXT-VQA (i.e., Conflicting Text) dataset, consisting of image-question pairs together with systematically generated persuasive prompts that deliberately conflict with visual evidence. Then, a thorough evaluation framework is designed and executed to benchmark the susceptibility of various models to these conflicting multimodal inputs. Comprehensive experiments over 11 state-of-the-art VLMs reveal that these models are indeed vulnerable to misleading textual prompts, often overriding clear visual evidence in favor of the conflicting text, and show an average performance drop of over 48.2% after only one round of persuasive conversation. Our findings highlight a critical limitation in current VLMs and underscore the need for improved robustness against textual manipulation.

视觉语言模型文本误导多模态鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。