测试视觉语言模型在城市中隐性需求下的探索能力,发现当前模型完成率仅21.1%。
CitySeeker: How Do VLMS Explore Embodied Urban Navigation With Implicit Human Needs?
- 构建城市隐性需求导航基准CitySeeker,包含6440条轨迹
- 顶级模型在长程推理中任务完成率仅为21.1%,主要受空间认知不足影响
- 提出人类认知启发的回溯、空间增强与记忆检索策略,助力智能体自主导航
视觉语言模型(VLMs)在显式指令导航上进展显著,但在动态城市环境中理解隐性人类需求(如“我口渴了”)的能力仍不充分。本文提出新基准CitySeeker,用于评估VLM在具身城市导航中应对隐性需求的空间推理与决策能力。该基准涵盖8个城市、6,440条轨迹,覆盖7种目标驱动场景和多样的视觉特征与隐性需求。大量实验表明,即使顶尖模型(如Qwen2.5-VL-32B-Instruct)任务完成率也仅有21.1%。分析发现主要瓶颈包括长时程推理中的错误累积、空间认知能力不足以及经验回忆缺失。为此,我们设计了受人类认知映射启发的探索策略——回溯机制、空间认知增强与基于记忆的检索(BCR),强调迭代观察-推理循环与路径自适应优化。该研究为发展具备鲁棒空间智能的VLM提供了可操作洞见,有助于解决“最后一公里”导航挑战。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have made significant progress in explicit instruction-based navigation; however, their ability to interpret implicit human needs (e.g., "I am thirsty") in dynamic urban environments remains underexplored. This paper introduces CitySeeker, a novel benchmark designed to assess VLMs' spatial reasoning and decision-making capabilities for exploring embodied urban navigation to address implicit needs. CitySeeker includes 6,440 trajectories across 8 cities, capturing diverse visual characteristics and implicit needs in 7 goal-driven scenarios. Extensive experiments reveal that even top-performing models (e.g., Qwen2.5-VL-32B-Instruct) achieve only 21.1% task completion. We find key bottlenecks in error accumulation in long-horizon reasoning, inadequate spatial cognition, and deficient experiential recall. To further analyze them, we investigate a series of exploratory strategies-Backtracking Mechanisms, Enriching Spatial Cognition, and Memory-Based Retrieval (BCR), inspired by human cognitive mapping's emphasis on iterative observation-reasoning cycles and adaptive path optimization. Our analysis provides actionable insights for developing VLMs with robust spatial intelligence required for tackling "last-mile" navigation challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。