arXiv:2603.18892cs.CVcs.AI2026-03被引 1

构建多跳空间推理基准,提升视觉语言模型的复杂空间理解能力。

MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model

  • 设计多跳复合空间推理任务,涵盖1至3跳复杂查询。
  • 提出Acc@50IoU指标,同时评估推理与精准定位能力。
  • 适合研究视觉-语言-动作系统与空间智能的学者使用。

空间推理是视觉语言模型(VLMs)的基础,尤其在作为物理环境中视觉-语言-动作(VLA)代理时至关重要。然而,现有基准大多聚焦于简单的单跳关系,忽视了真实场景中所需的多跳复合推理与精确视觉定位能力。为此,我们提出MultihopSpatial基准,包含三大贡献:(1) 一个面向多跳与复合空间推理的综合性基准,覆盖1至3跳的复杂查询及多样空间视角;(2) Acc@50IoU,一种结合答案选择与精确边界框预测的互补指标,评估推理与视觉定位双重能力,对稳健的VLA部署至关重要;(3) MultihopSpatial-Train,一个大规模训练语料库,用于培养空间智能。对37个先进VLMs的广泛评估揭示了八个关键发现,表明复合空间推理仍是重大挑战。最后,我们证明在该语料库上进行强化学习后训练,可显著提升VLM的内在空间推理能力及下游具身操作表现。

原文摘要 · Abstract (English)

Spatial reasoning is foundational for Vision-Language Models (VLMs), particularly when deployed as Vision-Language-Action (VLA) agents in physical environments. However, existing benchmarks predominantly focus on elementary, single-hop relations, neglecting the multi-hop compositional reasoning and precise visual grounding essential for real-world scenarios. To address this, we introduce MultihopSpatial, offering three key contributions: (1) A comprehensive benchmark designed for multi-hop and compositional spatial reasoning, featuring 1- to 3-hop complex queries across diverse spatial perspectives. (2) Acc@50IoU, a complementary metric that simultaneously evaluates reasoning and visual grounding by requiring both answer selection and precise bounding box prediction - capabilities vital for robust VLA deployment. (3) MultihopSpatial-Train, a dedicated large-scale training corpus to foster spatial intelligence. Extensive evaluation of 37 state-of-the-art VLMs yields eight key insights, revealing that compositional spatial reasoning remains a formidable challenge. Finally, we demonstrate that reinforcement learning post-training on our corpus enhances both intrinsic VLM spatial reasoning and downstream embodied manipulation performance.

空间推理视觉语言模型多跳推理具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。