arXiv:2510.19400cs.CV2025-10中稿 · ICLR被引 18

评测视觉语言模型在多视角机器人场景中的空间推理能力,发现当前模型远不及人类。

Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes

  • 构建多视角机器人基准测试集,评估模型跨视角空间理解能力。
  • 现有顶尖模型在多视角任务上表现远低于人类,存在显著差距。
  • 揭示空间智能与机器人执行能力正相关,单视角表现不等于多视角成功。

视觉语言模型(VLMs)是具身智能的核心,支撑机器人在复杂环境中的感知、推理与行动,也是近期视觉-语言-动作(VLA)模型的基础。然而,现有评估大多局限于单视角设置,忽视了多视角信息融合能力。随着多摄像头配置在机器人平台中日益普及,其可缓解遮挡与深度模糊问题。因此,VLM能否有效利用多视角输入进行机器人推理仍是一个开放问题。为此,我们提出MV-RoboBench,一个专为评估机器人操作场景中多视角空间推理能力设计的基准。该基准包含1.7k条人工标注的问答数据,涵盖八个子任务,分为空间理解与机器人执行两大类。我们评估了多种主流VLMs(开源与闭源),以及引入思维链启发技术的增强版本。结果表明,当前最先进模型仍远低于人类表现,凸显多视角机器人感知的巨大挑战。分析还发现:(i) 多视角场景下,空间智能与机器人任务执行能力呈正相关;(ii) 在通用单视角空间理解基准上的优异表现,并不能可靠转化为本基准中机器人任务的成功。我们已开源MV-RoboBench,提供数据与标准化评估协议,以推动具身视觉语言模型与VLA的发展。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are essential to Embodied AI, enabling robots to perceive, reason, and act in complex environments. They also serve as the foundation for the recent Vision-Language-Action (VLA) models. Yet most evaluations of VLMs focus on single-view settings, leaving their ability to integrate multi-view information underexplored. At the same time, multi-camera setups are increasingly standard in robotic platforms, as they provide complementary perspectives to mitigate occlusion and depth ambiguity. Whether VLMs can effectively leverage such multi-view inputs for robotic reasoning therefore remains an open question. To bridge this gap, we introduce MV-RoboBench, a benchmark specifically designed to evaluate the multi-view spatial reasoning capabilities of VLMs in robotic manipulation. MV-RoboBench consists of 1.7k manually curated QA items across eight subtasks, divided into two primary categories: spatial understanding and robotic execution. We evaluate a diverse set of existing VLMs, including both open-source and closed-source models, along with enhanced versions incorporating CoT-inspired techniques. The results show that state-of-the-art models remain far below human performance, underscoring the substantial challenges VLMs face in multi-view robotic perception. Additionally, our analysis uncovers two key findings: (i) spatial intelligence and robotic task execution are positively correlated in multi-view robotic scenarios; and (ii) strong performance on existing general-purpose single-view spatial understanding benchmarks does not reliably translate to success in the robotic spatial tasks assessed by our benchmark. We release MV-RoboBench as an open resource to foster progress in spatially grounded VLMs and VLAs, providing not only data but also a standardized evaluation protocol for multi-view embodied reasoning.

多视角空间推理机器人视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。