arXiv:2605.27315cs.CL2026-05

真实图像反而让视觉语言模型判断词汇具体性变差,尤其在图像无关时。

Real Images, Worse Judgments: Evaluating Vision-Language Models on Concreteness and Imagery

论文配图:Real Images, Worse Judgments: Evaluating Vision-Language Models on Concreteness and Imagery
图 1 · 摘自论文原文
  • 用人类对词汇具体性和意象的评分,测试模型是否能区分有用与无关图像信息。
  • 真实图像在低相关场景下使模型表现下降,对抽象词的判断准确率降低12%以上。
  • 仅依赖文本推理可显著改善模型在脆弱子集上的表现,适合需要精准语义理解的场景。

视觉输入常被视为提升多模态模型语言理解的助力。本文通过考察视觉-语言模型(VLMs)能否在词汇判断中区分有效视觉证据与偶然图像上下文,挑战这一假设。采用人类对词汇具体性和意象的评分作为基准,这些评分覆盖从抽象低意象词到具体高意象词的连续谱。实验发现,真实图像上下文并未带来一致收益,反而在视觉证据最不相关时显著损害模型与人类评分的一致性,尤其在抽象词上准确率下降超过12%。通过探针分析与典型相关分析,并结合归因案例研究,发现真实图像引发表征偏移,增强对虚假视觉线索的敏感性,导致目标词汇属性恢复能力减弱。进一步表明,在推理阶段仅依赖文本内容可缓解性能退化,尤其在脆弱子集上效果显著。研究提示当前指令微调的VLMs需更好校准视觉上下文介入时机。

原文摘要 · Abstract (English)

Visual inputs are often assumed to improve language understanding in multimodal models. We examine this assumption by asking whether vision-language models (VLMs) can distinguish useful visual evidence from incidental image context in lexical judgments. We use human concreteness and imagery ratings because they span words with varying expected visual relevance, from abstract and low-imagery words to concrete and high-imagery words. We find that real-image contexts do not yield consistent gains and often hurt alignment with human ratings, most sharply when visual evidence is least relevant. Through probing and canonical correlation analysis, complemented by an attribution case study, we find that real-image contexts are associated with representational shifts and greater sensitivity to spurious visual cues, coinciding with weaker recoverability of the targeted lexical properties. We further show that instructing models to focus solely on textual content at inference time can reduce this degradation, with the clearest gains on these vulnerable subsets. Our findings suggest that current instruction-tuned VLMs need better calibration of when visual context should inform lexical judgments.

多模态语义理解视觉偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。