arXiv:2604.07973cs.AI2026-04KDD

测试大模型在城市空中导航中的空间行动能力,发现距离人类水平仍有差距。

How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace

论文配图:How Far Are Large Multimodal Models from Human-Level Spatial Action? A Benchmark for Goal-Oriented Embodied Navigation in Urban Airspace
图 1 · 摘自论文原文
  • 构建5037个高精度城市三维导航样本,强调垂直动作与语义信息。
  • 现有大模型虽有初步行动能力,但错误在关键节点快速发散,无法持续导航。
  • 提出几何感知、跨视角理解等四条改进方向,适合具身智能研究者参考。

大型多模态模型(LMMs)在视觉-语言推理方面表现强劲,但在空间决策与行动能力上仍不明确。本文通过一个挑战性任务——城市三维空间中的目标导向导航,探究LMMs是否具备类人具身空间行动能力。我们耗时超500小时,构建了包含5,037个高质量导航样本的数据集,重点涵盖3D垂直动作和丰富的城市语义信息。随后对17种代表性模型(包括非推理型LMMs、推理型LMMs、基于智能体的方法及视觉-语言-动作模型)进行综合评估。实验表明,当前LMMs展现出初步的行动能力,但仍远未达到人类水平。此外,我们发现导航错误并非线性累积,而是在关键决策分叉点后迅速偏离目标。通过对这些关键点行为分析,揭示了模型局限性。最后,我们实证探索了四个潜在改进方向:几何感知、跨视角理解、空间想象与长期记忆。项目代码已开源:https://github.com/serenditipy-AC/Embodied-Navigation-Bench。

原文摘要 · Abstract (English)

Large multimodal models (LMMs) show strong visual-linguistic reasoning but their capacity for spatial decision-making and action remains unclear. In this work, we investigate whether LMMs can achieve embodied spatial action like human through a challenging scenario: goal-oriented navigation in urban 3D spaces. We first spend over 500 hours constructing a dataset comprising 5,037 high-quality goal-oriented navigation samples, with an emphasis on 3D vertical actions and rich urban semantic information. Then, we comprehensively assess 17 representative models, including non-reasoning LMMs, reasoning LMMs, agent-based methods, and vision-language-action models. Experiments show that current LMMs exhibit emerging action capabilities, yet remain far from human-level performance. Furthermore, we reveal an intriguing phenomenon: navigation errors do not accumulate linearly but instead diverge rapidly from the destination after a critical decision bifurcation. The limitations of LMMs are investigated by analyzing their behavior at these critical decision bifurcations. Finally, we experimentally explore four promising directions for improvement: geometric perception, cross-view understanding, spatial imagination, and long-term memory. The project is available at: https://github.com/serenditipy-AC/Embodied-Navigation-Bench.

具身智能多模态导航空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。