arXiv:2505.10453cs.CVcs.AI2025-05

测试视觉语言模型对虚拟物体的空间理解能力,发现其表现不佳。

Vision language models have difficulty recognizing virtual objects

  • 用未在图像中出现的虚拟物体测试模型场景理解能力
  • 主流视觉语言模型难以正确推理虚拟物体的空间关系
  • 适合研究多模态认知与场景理解的学者参考

视觉语言模型(VLMs)是结合语言和视觉编码器的多模态AI系统,可执行自动标注等复杂语义任务。然而,它们对图像所呈现场景的视觉空间属性的理解程度仍不明确。我们提出,通过描述虚拟物体——即图像中未实际出现的物体——来检验这些AI系统的场景理解能力。例如,一张展示一个人站在树下的图像,可搭配提示:想象一只风筝卡在树上。能够真正理解场景的VLM应能更新其表征,并合理推断三个物体之间的空间关系。我们对当前最先进的VLMs进行了系统性评估,结果表明,它们处理虚拟物体的能力严重不足。

原文摘要 · Abstract (English)

Vision language models (VLMs) are AI systems paired with both language and vision encoders to process multimodal input. They are capable of performing complex semantic tasks such as automatic captioning, but it remains an open question about how well they comprehend the visuospatial properties of scenes depicted in the images they process. We argue that descriptions of virtual objects -- objects that are not visually represented in an image -- can help test scene comprehension in these AI systems. For example, an image that depicts a person standing under a tree can be paired with the following prompt: imagine that a kite is stuck in the tree. VLMs that comprehend the scene should update their representations and reason sensibly about the spatial relations between all three objects. We describe systematic evaluations of state-of-the-art VLMs and show that their ability to process virtual objects is inadequate.

多模态理解视觉语言模型空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。