测试视觉语言模型跨视角定位能力,发现其在第三人称视角下表现明显下降。
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
- 构建首个多视角空间定位评测基准ViewSpatial-Bench,支持五类任务。
- 模型在摄像头视角准确率高,换到人视角后性能显著下降。
- 通过多视角微调提升46.24%整体表现,适合研究具身智能的空间理解。
视觉语言模型(VLMs)在理解与推理视觉内容方面表现出色,但在需要跨视角理解与空间推理的任务中仍面临重大挑战。我们识别出一个关键局限:当前的VLMs主要擅长以摄像机为中心的空间推理,却无法有效迁移到需采用其他实体空间参照系的分配视角。为此,我们提出了ViewSpatial-Bench,首个专为多视角空间定位识别设计的综合性基准,涵盖五种不同任务类型,并依托自动化3D标注流程生成精确的方向标签。对多种VLMs在ViewSpatial-Bench上的全面评估揭示了显著的性能差距:模型在摄像机视角任务中表现良好,但在人类视角推理时准确率明显降低。通过在我们的多视角空间数据集上微调VLMs,我们在各类任务上实现了46.24%的整体性能提升,验证了该方法的有效性。本工作为具身AI系统中的空间智能建立了重要基准,并提供了实证证据,表明建模三维空间关系能显著增强VLMs的空间理解能力。
原文摘要 · Abstract (English)
Vision-language models (VLMs) have demonstrated remarkable capabilities in understanding and reasoning about visual content, but significant challenges persist in tasks requiring cross-viewpoint understanding and spatial reasoning. We identify a critical limitation: current VLMs excel primarily at egocentric spatial reasoning (from the camera's perspective) but fail to generalize to allocentric viewpoints when required to adopt another entity's spatial frame of reference. We introduce ViewSpatial-Bench, the first comprehensive benchmark designed specifically for multi-viewpoint spatial localization recognition evaluation across five distinct task types, supported by an automated 3D annotation pipeline that generates precise directional labels. Comprehensive evaluation of diverse VLMs on ViewSpatial-Bench reveals a significant performance disparity: models demonstrate reasonable performance on camera-perspective tasks but exhibit reduced accuracy when reasoning from a human viewpoint. By fine-tuning VLMs on our multi-perspective spatial dataset, we achieve an overall performance improvement of 46.24% across tasks, highlighting the efficacy of our approach. Our work establishes a crucial benchmark for spatial intelligence in embodied AI systems and provides empirical evidence that modeling 3D spatial relationships enhances VLMs' corresponding spatial comprehension capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。