测试视觉语言模型如何理解空间方向,发现它们在跨语言中表现不稳定。
Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities
- 构建COMFORT评估框架,系统测试模型对空间参照系的推理能力。
- 模型在多语言测试中表现不一致,英文主导导致其他语言偏差。
- 揭示模型缺乏文化多样性适应力,提示需关注人类认知的复杂性。
情景化交流中的空间表达常具歧义,其含义依赖于说话者与听者的空间参照系(FoR)。尽管视觉语言模型(VLMs)在空间语言理解与推理方面备受关注,但其潜在歧义仍鲜有研究。为此,我们提出COnsistent Multilingual Frame Of Reference Test(COMFORT),一个系统评估VLM空间推理能力的评测协议。使用COMFORT评估九个主流VLMs,结果表明:尽管模型在解析英语惯例时有一定对齐,但存在显著缺陷——(1)鲁棒性与一致性差;(2)无法灵活适应多种参照系;(3)跨语言测试中未能遵循语言或文化特异性规范,英语主导现象明显。随着视觉语言模型向人类认知对齐的努力推进,亟需关注空间推理的模糊性与跨文化多样性。
原文摘要 · Abstract (English)
Spatial expressions in situated communication can be ambiguous, as their meanings vary depending on the frames of reference (FoR) adopted by speakers and listeners. While spatial language understanding and reasoning by vision-language models (VLMs) have gained increasing attention, potential ambiguities in these models are still under-explored. To address this issue, we present the COnsistent Multilingual Frame Of Reference Test (COMFORT), an evaluation protocol to systematically assess the spatial reasoning capabilities of VLMs. We evaluate nine state-of-the-art VLMs using COMFORT. Despite showing some alignment with English conventions in resolving ambiguities, our experiments reveal significant shortcomings of VLMs: notably, the models (1) exhibit poor robustness and consistency, (2) lack the flexibility to accommodate multiple FoRs, and (3) fail to adhere to language-specific or culture-specific conventions in cross-lingual tests, as English tends to dominate other languages. With a growing effort to align vision-language models with human cognitive intuitions, we call for more attention to the ambiguous nature and cross-cultural diversity of spatial reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。