测试视觉语言模型在四种语言中使用空间指示词的能力
Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models
- 构建多语言空间指示词评测基准,考察模型对上下文依赖指代的理解
- 模型在根据物体距离选择合适指示词上表现不如人类,存在系统性偏差
- 适合关注多语言视觉推理与跨文化语义差异的研究者
视觉语言模型(VLMs)应具备基于文本和图像进行空间推理的能力。为评估这一能力,本文聚焦空间指示表达(spatial deictic expressions),即其指代对象由情境决定的表达,如“this”和“that”。处理这类表达需模型联合推理语言与视觉空间,将上下文相关的指代准确锚定于图像的空间结构中。同时,跨语言选择恰当的指示词还需理解各语言特有的空间区分机制。本文构建了一个涵盖四种语言的评测基准,用于评估VLMs在使用空间指示词方面的多语言能力。实验结果表明,所测模型在使用指示词时的行为与人类存在显著差异,尤其在依据物体距离选择合适指示词方面表现不佳。
原文摘要 · Abstract (English)
One of the expected abilities of vision-language models (VLMs) is spatial reasoning ability based on a given text and image. To evaluate the spatial reasoning abilities of VLMs, we focus on the use of spatial deictic expressions, which are defined as spatial expressions whose referent is determined by their situational context, such as ``this'' and ``that''. To handle spatial deictic expressions, VLMs must jointly reason over language and visual space, grounding context-dependent references in the image's spatial structure. In addition, selecting appropriate spatial deictic expressions across languages requires VLMs to understand the language-specific spatial distinctions encoded by these expressions. In this paper, we develop a benchmark to evaluate the multilingual ability of VLMs to use spatial deictic expressions in four languages. Our experiments using this benchmark reveal that the tested models use demonstratives in a manner different from that of humans, particularly in selecting the appropriate demonstratives based on the distance to the object.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。