测试视觉语言模型在3D场景中从空间推理到行动的转化能力
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes

- 构建仿真环境下的多轮交互基准,评估模型空间推理转行动力
- 模型在单任务上表现良好,多轮反馈中空间认知持续性差
- 适合关注具身智能、空间理解与多轮决策的研究者
人类能自然感知三维空间布局,形成认知表征,进行空间关系推理,并将其转化为行动。尽管近期视觉语言模型(VLMs)在基于观察的空间感知与推理任务上表现良好,但它们是否能建立连贯的空间理解、据此采取行动并利用多轮反馈优化行为仍不清楚。为此,我们提出 extbf{SpatialAct},一个基于模拟器的基准,用于探测3D场景中的动作条件空间推理能力。从最具挑战性的多轮交互优化设置出发,我们进一步设计其分解版本——单步错误检测与修正,并构建五个基础空间能力任务,以诊断模型失败的根本原因。实验表明存在明显的推理-行动差距:当前VLMs在孤立空间推理任务上表现优异,但在多轮反馈过程中难以维持连贯的空间信念,导致行动不可靠,显著落后于人类。结果表明,即使低层控制被抽象化,当前VLM代理在动作引发环境变化下仍缺乏稳健的空间状态追踪能力。
原文摘要 · Abstract (English)
Humans can effortlessly perceive spatial layouts, form cognitive representations, reason about spatial relations, and translate such reasoning into actions in everyday 3D environments. Although recent vision-language models (VLMs) have shown promising performance on observation-conditioned spatial perception and reasoning tasks, it remains unclear whether they can build coherent spatial understanding, act upon it, and refine their actions through multi-turn feedback. To study this problem, we introduce \textbf{SpatialAct}, a simulator-grounded benchmark for probing \textit{action-conditioned spatial reasoning} in 3D scenes. Starting from the most challenging setting, Multi-turn Interactive Refinement, we further design its decomposed counterpart, Single-step Error Detection and Fix, together with five fundamental spatial ability tasks to diagnose the underlying causes of model failures. Experiments reveal a clear reasoning-to-action gap: current VLMs can perform well on isolated spatial reasoning tasks, but struggle to maintain coherent spatial beliefs and produce reliable actions during multi-turn feedback, substantially underperforming humans. These results suggest that current VLM agents still lack robust spatial state tracking under action-induced environment changes, even when low-level control is abstracted away.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。