评测视觉语言模型对连续空间的感知能力,发现多数模型存在短板。
CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models
- 构建多图像连续空间理解基准CoSpace,聚焦静态视角下的空间连贯性。
- 在19个模型上测试,发现多数模型在连续空间感知上表现不佳。
- 开源与闭源模型差异在于响应一致性,而非准确率高低。
视觉语言模型(VLMs)在视觉理解方面取得显著进展。随着图像上下文长度增加,模型能理解更广泛的视角和空间。现有基准多关注非空间相关图像或不同视角的离散图像,忽视了从固定视角观察时产生的空间连续图像的组合特性。我们称之为连续空间感知(Continuous Space Perception)。当从固定视角改变方向时,生成一系列空间连续的图像,可重构完整空间。本文提出CoSpace,一个用于评估VLM连续空间感知能力的多图像理解基准,包含2,918张图像和1,626个问答对,涵盖七类任务。我们在19个专有及开源VLM上进行评估,结果表明大多数模型在连续空间感知方面存在缺陷,包括部分专有模型。有趣的是,开源与专有模型的主要差异不在于准确性,而在于响应的一致性。我们认为提升连续空间感知能力对VLM在真实场景中的有效应用至关重要,呼吁进一步研究。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have recently witnessed significant progress in visual comprehension. As the permitting length of image context grows, VLMs can now comprehend a broader range of views and spaces. Current benchmarks provide insightful analysis of VLMs in tasks involving complex visual instructions following, multi-image understanding and spatial reasoning. However, they usually focus on spatially irrelevant images or discrete images captured from varied viewpoints. The compositional characteristic of images captured from a static viewpoint remains underestimated. We term this characteristic as Continuous Space Perception. When observing a scene from a static viewpoint while shifting orientations, it produces a series of spatially continuous images, enabling the reconstruction of the entire space. In this paper, we present CoSpace, a multi-image visual understanding benchmark designed to assess the Continuous Space perception ability for VLMs. CoSpace contains 2,918 images and 1,626 question-answer pairs, covering seven types of tasks. We conduct evaluation across 19 proprietary and open-source VLMs. Results reveal that there exist pitfalls on the continuous space perception ability for most of the evaluated models, including proprietary ones. Interestingly, we find that the main discrepancy between open-source and proprietary models lies not in accuracy but in the consistency of responses. We believe that enhancing the ability of continuous space perception is essential for VLMs to perform effectively in real-world tasks and encourage further research to advance this capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。