arXiv:2502.04359cs.CLcs.AI2025-02IJCAI被引 3

用指称表达任务评估视觉语言模型的空间推理能力

Exploring Spatial Language Grounding Through Referring Expressions

  • 以指称表达理解任务为平台,分析模型空间推理能力
  • 在模糊检测、复杂句式和否定表达下,模型表现显著下降
  • 揭示不同模型在拓扑、方向、距离等语义上的差异表现

空间推理是人类认知的重要组成部分,也是当前视觉-语言模型(VLMs)面临挑战的领域。现有分析多依赖图像描述和视觉问答任务,本文提出改用指称表达理解(Referring Expression Comprehension)任务作为评估VLMs空间推理能力的新平台。该平台能深入分析模型在三种关键情境下的表现:1)目标检测存在歧义;2)表达包含长句结构与多重空间关系;3)涉及否定语义(如'not')。我们采用任务专用架构及大型VLMs进行实验,揭示其在上述情境中的优劣势。结果表明,尽管所有模型均面临困难,但其相对表现取决于底层架构及具体空间语义类别(如拓扑、方向、邻近等)。研究揭示了现有方法的局限性,并指明未来研究方向。

原文摘要 · Abstract (English)

Spatial Reasoning is an important component of human cognition and is an area in which the latest Vision-language models (VLMs) show signs of difficulty. The current analysis works use image captioning tasks and visual question answering. In this work, we propose using the Referring Expression Comprehension task instead as a platform for the evaluation of spatial reasoning by VLMs. This platform provides the opportunity for a deeper analysis of spatial comprehension and grounding abilities when there is 1) ambiguity in object detection, 2) complex spatial expressions with a longer sentence structure and multiple spatial relations, and 3) expressions with negation ('not'). In our analysis, we use task-specific architectures as well as large VLMs and highlight their strengths and weaknesses in dealing with these specific situations. While all these models face challenges with the task at hand, the relative behaviors depend on the underlying models and the specific categories of spatial semantics (topological, directional, proximal, etc.). Our results highlight these challenges and behaviors and provide insight into research gaps and future directions.

空间推理视觉语言指称表达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。