构建仿真环境诊断视觉语言模型空间推理能力,助力机器人任务落地。
ESPIRE: A Diagnostic Benchmark for Embodied Spatial Reasoning of Vision-Language Models
- 将任务拆解为定位与执行,以生成式方式评估空间推理
- 在仿真环境中测试多个机器人任务,覆盖广泛的空间推理场景
- 可细粒度分析模型从理解到行动的推理过程,适合研究者优化模型
近期视觉语言模型(VLMs)致力于提升具身领域的空间认知能力。然而,现有评估方法在范式和覆盖范围上均存在局限,制约了模型的快速迭代。为此,我们提出ESPIRE——一个面向具身空间推理的诊断基准。ESPIRE构建了一个物理化仿真实验环境,使VLMs在具身任务中接受空间推理评估,缩小了评测与实际部署之间的差距。为适配机器人任务,我们将每个任务分解为定位与执行两个阶段,并将其均建模为生成问题,与主流依赖干扰项、忽略执行过程的判别式评估(如视觉问答)形成鲜明对比。该分解支持从被动推理到决策行动的细粒度分析。我们在指令层与环境层系统设计了ESPIRE,确保涵盖多样化的空间推理场景。利用该基准,我们诊断了一系列前沿VLMs,并深入分析其空间推理行为。
原文摘要 · Abstract (English)
A recent trend in vision-language models (VLMs) has been to enhance their spatial cognition for embodied domains. Despite progress, existing evaluations have been limited both in paradigm and in coverage, hindering rapid, iterative model development. To address these limitations, we propose ESPIRE, a diagnostic benchmark for embodied spatial reasoning. ESPIRE offers a simulated world that physically grounds VLMs and evaluates them on spatial-reasoning-centric robotic tasks, thus narrowing the gap between evaluation and real-world deployment. To adapt VLMs to robotic tasks, we decompose each task into localization and execution, and frame both as generative problems, in stark contrast to predominant discriminative evaluations (e.g., via visual-question answering) that rely on distractors and discard execution. This decomposition further enables a fine-grained analysis beyond passive spatial reasoning toward reasoning to act. We systematically design ESPIRE both at the instruction level and at the environment level, ensuring broad coverage of spatial reasoning scenarios. We use ESPIRE to diagnose a range of frontier VLMs and provide in-depth analysis of their spatial reasoning behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。