arXiv:2606.09669cs.AIcs.CL2026-06被引 2

构建真实世界任务的交互式空间推理评测基准,检验多模态模型在复杂场景中的理解与操作能力。

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

论文配图:SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
图 1 · 摘自论文原文
  • 统一协议整合八种仿真环境,支持多模态智能体在真实任务中主动探索与决策。
  • 15个先进模型平均成功率仅17.4%,最强模型仍难突破20%,暴露长期规划短板。
  • 适合研究多模态推理、具身智能与空间认知的学者,尤其关注实际交互性能的团队。

空间推理是多模态大语言模型感知与操作物理世界的基础能力。然而,现有评测主要依赖被动评估(如静态视觉问答)或特定仿真器流程,难以衡量通用的交互式空间理解能力。我们提出 SpatialWorld,一个专为评估多模态智能体在复杂真实任务中交互式空间理解而设计的统一基准。该基准在共享的仿真无关协议下集成八种异构仿真后端,包含760个由人类标注的任务,覆盖家庭日常、出行、社交协作等多样领域。智能体需在仅视觉输入的部分可观测条件下,主动收集第一人称视觉证据,并通过适配多模态大语言模型的文本动作接口表达决策。为保证评估可靠性,每项任务均配备人工验证的初始状态、参考执行轨迹与终端状态验证器。对15个先进模型的评估显示,稳健的空间任务求解依然困难:最强模型GPT-5的平均任务成功率(TSR)仅为17.4%,领先开源模型Qwen-3.5达到14.1%。进一步分析揭示任务成功率与执行效率之间存在明显差距,且不同领域表现差异显著。这些在主动探索与长时程规划方面的瓶颈,使SpatialWorld成为未来空间智能体研发的严格测试平台。

原文摘要 · Abstract (English)

Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world. However, existing benchmarks predominantly rely on passive evaluation (e.g., static VQA) or simulator-specific pipelines, failing to assess general interactive spatial understanding. We introduce SpatialWorld, a unified benchmark designed specifically for evaluating the interactive spatial understanding of multimodal agents in complex real-world tasks. Integrating eight heterogeneous simulation backends under a shared, simulator-agnostic protocol, SpatialWorld features 760 human-annotated tasks across diverse domains (e.g., household routines, travel, social collaboration). Agents must solve tasks under vision-only partial observability, actively gathering egocentric visual evidence and expressing decisions via a unified, text-based action interface native to MLLMs. For reliable evaluation, each task includes a human-validated initial state, a reference trajectory, and a terminal-state verifier. Evaluating 15 advanced agents reveals that robust spatial task solving remains challenging: the strongest model, GPT-5, achieves an average task success rate (TSR) of only 17.4%, while the leading open-source model, Qwen-3.5, reaches 14.1%. Further analysis exposes a clear mismatch between task success and execution efficiency, alongside substantial domain-specific performance variations. These bottlenecks in active exploration and long-horizon planning position SpatialWorld as a rigorous testbed for future spatial agents.

空间推理多模态具身智能评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。