arXiv:2512.15933cs.CV2025-12中稿 · EACL 2026被引 2

用大规模视觉导航测试多模态大模型真实城市导航能力

City Navigation in the Wild: Exploring Emergent Navigation from Web-Scale Knowledge in MLLMs

  • 让模型仅凭视觉和内部推理,在真实城市中连续决策50多个节点
  • 现有顶尖模型在无标注环境下导航成功率不足30%
  • 提出路径显式化方法,通过提取城市认知地图显著提升性能

利用多模态大语言模型(MLLMs)构建具身智能体,有望解决复杂现实任务。然而,当前评估基准仍以语言为中心或严重依赖模拟环境,难以检验真实世界所需的知识密集型推理能力。为此,我们提出稀疏接地视觉导航任务,专门评估MLLM在挑战性、知识密集型真实环境中的序列决策能力。我们构建了涵盖四座全球城市的CityNav基准,用于评估纯MLLM驱动的智能体在城市导航中的表现。智能体需仅依靠视觉输入和内部多模态推理,完成超过50个决策点的导航,无需额外环境标注或特殊架构调整。关键在于,智能体必须自主通过解读城市特征线索和识别地标实现定位,进行空间推理,并策略性规划与执行路线。大量实验表明,当前最先进MLLMs(如GEPA、思维链、反思机制)及竞争基线PReP在此设定下表现不佳。为此,我们提出路径显式化(VoP),通过从MLLM中探查城市级认知地图(关键地标与目的地方向),显式化内部推理,显著提升导航成功率。

原文摘要 · Abstract (English)

Leveraging multimodal large language models (MLLMs) to develop embodied agents offers significant promise for addressing complex real-world tasks. However, current evaluation benchmarks remain predominantly language-centric or heavily reliant on simulated environments, rarely probing the nuanced, knowledge-intensive reasoning essential for practical, real-world scenarios. To bridge this critical gap, we introduce the task of Sparsely Grounded Visual Navigation, explicitly designed to evaluate the sequential decision-making abilities of MLLMs in challenging, knowledge-intensive real-world environment. We operationalize this task with CityNav, a comprehensive benchmark encompassing four diverse global cities, specifically constructed to assess raw MLLM-driven agents in city navigation. Agents are required to rely solely on visual inputs and internal multimodal reasoning to sequentially navigate 50+ decision points without additional environmental annotations or specialized architectural modifications. Crucially, agents must autonomously achieve localization through interpreting city-specific cues and recognizing landmarks, perform spatial reasoning, and strategically plan and execute routes to their destinations. Through extensive evaluations, we demonstrate that current state-of-the-art MLLMs, reasoning techniques (e.g., GEPA, chain-of-thought, reflection) and competitive baseline PReP significantly underperform in this challenging setting. To address this, we propose Verbalization of Path(VoP), which explicitly grounds the agent's internal reasoning by probing city-scale cognitive maps (key landmarks and directions toward the destination) from the MLLM, substantially enhancing navigation success. Project Webpage: https://dwipddalal.github.io/AgentNav/

城市导航多模态大模型具身智能认知地图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。