arXiv:2608.27456cs.CV2026-08

测试大模型在真实城市中长期探索的能力,发现其局部感知难转为持续导航。

UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City

论文配图:UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City
图 1 · 摘自论文原文
  • 构建香港真实尺度3D城市沙盒,支持第一人称闭环交互
  • 模型能短期识别场景但远距离导航与路径变化时迅速失效
  • 揭示当前大模型在开放城市环境中的探索局限性,适合研究智能体规划

多模态大语言模型(MLLM)可解析街景图像,但城市智能体的行动能力取决于局部感知在移动后是否仍有效。本文通过提出UrbanGround——首个基于全港3D地理空间数据构建的物理约束真实城市模拟环境,首次使该问题可测试。该平台支持第一人称视角的闭环交互,并提供可导航的交互地图。我们从三个层面分析空间认知演进:首先检验智能体能否在主动观察后准确回答空间问题;其次考察其导航能力在目标距离增加、提示不明确时的表现;最后评估其在路线变更与行人运动变化下的行为鲁棒性。实验表明,当前MLLM智能体虽具备良好的视觉识别与短程空间推理能力,但方向判断与行人避让仍不可靠;在长距离探索中,局部能力无法形成持续目标导向行为,错误累积且缺乏有效修正。我们希望UrbanGround能推动对当前模型在复杂开放城市环境中可靠探索边界的深入研究。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) can interpret a street view, but urban agency depends on whether such local evidence remains useful after the agent starts to move. In this paper, we investigate how far current MLLM agents can turn local urban perception into reliable action in a complicated real-scale city. We propose UrbanGround, the first sandbox to make this question testable in a physically constrained replica of Hong Kong built from territory-wide 3D geospatial data. UrbanGround supports closed-loop interaction from a first-person view and provides an interactive map for navigation. Agents can directly enter the 3D city and explore from a first-person view. Our analysis follows the growth of the spatial problem through three research questions. We first test whether an agent can ground a local scene well enough to answer spatial questions after active observation. Then we ask whether that grounding supports navigation as destinations become farther away and less explicit. Finally, we examine whether the resulting behavior survives changes in route availability and pedestrian motion. Contemporary MLLM agents usually show useful atomic abilities in visual recognition and short-range spatial reasoning, while orientation and pedestrian-aware movement remain unreliable. Their central failure emerges over extended exploration, where local abilities do not compose into sustained goal-directed behavior and errors accumulate without effective correction. We hope UrbanGround will support broader study of how far current MLLM agents can explore reliably in complex, open-ended urban environments.

城市智能体多模态模型长程规划仿真环境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。