用指称表达任务检测视觉语言模型的空间理解能力
Referring Expressions as a Lens into Spatial Language Grounding in Vision-Language Models
- 以指称表达理解任务评估模型空间推理能力
- 复杂空间描述和否定句让模型表现显著下降
- 揭示不同模型在拓扑、方向等语义上的差异
空间推理是人类认知的重要组成部分,但最新视觉语言模型(VLMs)在此方面表现不佳。现有分析多依赖图像字幕和视觉问答任务,本文提出改用指称表达理解(Referring Expression Comprehension)任务作为评估平台,以更深入分析模型在物体检测模糊、长句含多重空间关系、以及含否定词(如'not')情况下的空间理解与定位能力。我们采用任务专用架构和大型VLMs进行实验,发现所有模型在该任务中均面临挑战,其相对表现取决于底层模型结构及具体空间语义类别(如拓扑、方向、邻近等)。结果揭示了当前模型的局限性,为后续研究提供了方向。
原文摘要 · Abstract (English)
Spatial Reasoning is an important component of human cognition and is an area in which the latest Vision-language models (VLMs) show signs of difficulty. The current analysis works use image captioning tasks and visual question answering. In this work, we propose using the Referring Expression Comprehension task instead as a platform for the evaluation of spatial reasoning by VLMs. This platform provides the opportunity for a deeper analysis of spatial comprehension and grounding abilities when there is 1) ambiguity in object detection, 2) complex spatial expressions with a longer sentence structure and multiple spatial relations, and 3) expressions with negation ('not'). In our analysis, we use task-specific architectures as well as large VLMs and highlight their strengths and weaknesses in dealing with these specific situations. While all these models face challenges with the task at hand, the relative behaviors depend on the underlying models and the specific categories of spatial semantics (topological, directional, proximal, etc.). Our results highlight these challenges and behaviors and provide insight into research gaps and future directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。