arXiv:2506.05146cs.CVcs.CL2025-06EMNLP被引 6

测试视觉语言模型对物体属性与关系的理解能力,发现其表现远低于人类。

CIVET: Systematic Evaluation of Understanding in VLMs

  • 设计可控刺激框架CIVET,系统评估模型对物体属性与关系的理解。
  • 仅少数基础属性识别准确,且性能受物体位置影响显著。
  • 在理解物体间基本关系上表现差,仍不及人类水平。

尽管视觉语言模型(VLMs)在多个任务中表现出色,但其对场景底层结构与语义的理解仍缺乏系统研究。为此,我们提出CIVET——一种基于可控刺激的系统性评估框架,用于探究模型对物体属性与关系的理解能力。该框架克服了传统评估中缺乏标准化、存在标注噪声与数据集偏差的问题。利用CIVET,我们在无噪声、控制复杂的刺激集上评估了五种主流VLMs。结果表明:1)当前模型仅能准确识别有限的基本物体属性;2)性能高度依赖物体在场景中的位置;3)在理解物体间基本关系方面表现不佳。与人类标注者对比显示,现有VLMs尚未达到人类级理解水平。

原文摘要 · Abstract (English)

While Vision-Language Models (VLMs) have achieved competitive performance in various tasks, their comprehension of the underlying structure and semantics of a scene remains understudied. To investigate the understanding of VLMs, we study their capability regarding object properties and relations in a controlled and interpretable manner. To this scope, we introduce CIVET, a novel and extensible framework for systematiC evaluatIon Via controllEd sTimuli. CIVET addresses the lack of standardized systematic evaluation for assessing VLMs' understanding, enabling researchers to test hypotheses with statistical rigor. With CIVET, we evaluate five state-of-the-art VLMs on exhaustive sets of stimuli, free from annotation noise, dataset-specific biases, and uncontrolled scene complexity. Our findings reveal that 1) current VLMs can accurately recognize only a limited set of basic object properties; 2) their performance heavily depends on the position of the object in the scene; 3) they struggle to understand basic relations among objects. Furthermore, a comparative evaluation with human annotators reveals that VLMs still fall short of achieving human-level accuracy.

视觉语言模型理解评估可控实验属性识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。