arXiv:2509.22437cs.CLcs.AI2025-09被引 5

提出Chimera评测集,揭示视觉语言模型看图答题的三大作弊行为。

Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding

  • 构建7500张维基百科图表数据集,含语义三元组与多层级问题
  • 发现模型90%以上高分源于'聪明汉斯'等表面捷径而非真实理解
  • 适用于评估模型是否真懂图表,适合关注AI可解释性的研究者

图表以可视化方式呈现符号信息,对人工智能模型构成挑战。尽管近期评估显示视觉语言模型(VLMs)在图表基准上表现良好,但其依赖知识、推理或模态捷径的问题令人担忧。为此,我们提出Chimera,一个包含7500张高质量维基百科图表的综合性评测集,每张图表均标注语义三元组,并配有针对实体识别、关系理解、知识定位和视觉推理四个核心维度的多层级问题。通过Chimera,我们检测到三种常见答题捷径:(1)视觉记忆捷径,模型依赖已记住的视觉模式;(2)知识回忆捷径,模型依赖预存事实而非解读图表;(3)聪明汉斯捷径,模型利用表面语言模式或先验知识。我们在15个开源VLMs(来自7个模型家族)上进行测试,发现其看似优异的表现主要源于捷径行为:视觉记忆影响较小,知识回忆中等,而聪明汉斯捷径贡献显著。该结果暴露当前VLM在复杂视觉输入(如图表)理解上的关键缺陷,强调需建立更鲁棒的评估体系,真正衡量模型对视觉内容的理解能力。

原文摘要 · Abstract (English)

Diagrams convey symbolic information in a visual format rather than a linear stream of words, making them especially challenging for AI models to process. While recent evaluations suggest that vision-language models (VLMs) perform well on diagram-related benchmarks, their reliance on knowledge, reasoning, or modality shortcuts raises concerns about whether they genuinely understand and reason over diagrams. To address this gap, we introduce Chimera, a comprehensive test suite comprising 7,500 high-quality diagrams sourced from Wikipedia; each diagram is annotated with its symbolic content represented by semantic triples along with multi-level questions designed to assess four fundamental aspects of diagram comprehension: entity recognition, relation understanding, knowledge grounding, and visual reasoning. We use Chimera to measure the presence of three types of shortcuts in visual question answering: (1) the visual-memorization shortcut, where VLMs rely on memorized visual patterns; (2) the knowledge-recall shortcut, where models leverage memorized factual knowledge instead of interpreting the diagram; and (3) the Clever-Hans shortcut, where models exploit superficial language patterns or priors without true comprehension. We evaluate 15 open-source VLMs from 7 model families on Chimera and find that their seemingly strong performance largely stems from shortcut behaviors: visual-memorization shortcuts have slight impact, knowledge-recall shortcuts play a moderate role, and Clever-Hans shortcuts contribute significantly. These findings expose critical limitations in current VLMs and underscore the need for more robust evaluation protocols that benchmark genuine comprehension of complex visual inputs (e.g., diagrams) rather than question-answering shortcuts.

视觉语言模型图表理解评测集捷径学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。