arXiv:2504.16061cs.CVcs.AI2025-04被引 6

VLM在简单空间认知任务中表现不可靠,微小提示差异会导致结果波动。

Vision language models are unreliable at trivial spatial cognition

  • 构建TableTest数据集,测试视觉语言模型对物体相对位置的判断能力。
  • 相同场景下,逻辑等价的提示导致模型性能显著下降。
  • 揭示VLM在真实应用中处理空间关系的局限性,适合关注模型鲁棒性的研究者。

视觉语言模型(VLMs)旨在从图像中提取相关视觉空间信息。尽管有研究认为VLMs具备类人场景理解能力,但也有证据显示其在处理关系信息时存在困难。为实现广泛应用,VLMs必须在多种相关任务中表现稳定可靠。本文测试了这些架构在基础空间认知任务中的可靠性,例如在无杂乱场景中识别一个物体是否位于另一个物体左侧。为此,我们构建了一个名为TableTest的基准数据集,其图像描绘了3D场景中物体在桌面上的排列,并用以评估最先进的VLMs。结果显示,使用逻辑等价描述的微小提示变化可导致性能下降。分析表明,VLM在现实应用中推理空间关系存在局限性。同时,这也为改进图像字幕语料库以实现更高效训练与测试提供了新机遇。

原文摘要 · Abstract (English)

Vision language models (VLMs) are designed to extract relevant visuospatial information from images. Some research suggests that VLMs can exhibit humanlike scene understanding, while other investigations reveal difficulties in their ability to process relational information. To achieve widespread applicability, VLMs must perform reliably, yielding comparable competence across a wide variety of related tasks. We sought to test how reliable these architectures are at engaging in trivial spatial cognition, e.g., recognizing whether one object is left of another in an uncluttered scene. We developed a benchmark dataset -- TableTest -- whose images depict 3D scenes of objects arranged on a table, and used it to evaluate state-of-the-art VLMs. Results show that performance could be degraded by minor variations of prompts that use logically equivalent descriptions. These analyses suggest limitations in how VLMs may reason about spatial relations in real-world applications. They also reveal novel opportunities for bolstering image caption corpora for more efficient training and testing.

视觉语言模型空间认知模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。