arXiv:2606.08992cs.ROcs.AI2026-06被引 1

提出空间认知记忆与任务引导空间推理,实现零样本环境导航新突破

SpaceVLN: A Zero-Shot Vision-and-Language Navigation Agent with Online Spatial Cognitive Memory and Reasoning

论文配图:SpaceVLN: A Zero-Shot Vision-and-Language Navigation Agent with Online Spatial Cognitive Memory and Reasoning
图 1 · 摘自论文原文
  • 构建分阶段闭环框架,动态维护空间路标与层级记忆
  • 在多个数据集上达顶尖零样本性能,支持复杂路径规划
  • 适用于真实机器人部署,适合需要泛化能力的导航研究者

连续环境中的视觉语言导航要求智能体理解未见过环境的空间结构以遵循语言指令。尽管基础模型为无特定任务策略训练的零样本导航开辟了新路径,但许多导航器仍依赖局部视觉线索和线性历史推理,忽视了探索区域、行进路径、地标及其空间关系的空间特性。本文提出SpaceVLN,一个围绕空间认知记忆与任务引导空间推理构建的导航智能体。具体而言,SpaceVLN引入高效的分阶段闭环框架,将规划与执行组织为可验证的空间-地标阶段。导航过程中,智能体逐步将已探索区域抽象为空间路标,并动态维护子任务相关的地标证据,形成用于定位与空间关系理解的层次化空间认知记忆。基于该记忆,空间思维链(Spatial-CoT)整合任务进展推理与空间感知、分析、预测,实现任务引导的空间推理,支持具身导航。统一的阶段接口使SpaceVLN能在无需特定任务策略训练的情况下,同时处理视觉语言导航与目标物体导航。在R2R-CE、RxR-CE、GN-Bench和HM3D-OVON多个数据集上,SpaceVLN达到当前最优的零样本性能,真实机器人部署进一步验证其适用性。结果表明,空间认知记忆与任务引导空间推理是构建更强具身导航智能体的实用基础。

原文摘要 · Abstract (English)

Vision-and-Language Navigation in continuous environments requires agents to understand the spatial structure of previously unseen environments in order to follow language instructions. Although foundation models have opened a promising path toward zero-shot navigation without task-specific policy training, many navigators still rely on local visual cues and linear history-based reasoning, overlooking the spatial nature of navigation across explored regions, traversed paths, landmarks, and their spatial relations. In this paper, we propose SpaceVLN, a navigation agent built around Spatial Cognitive Memory and Task-Guided Spatial Reasoning. Specifically, SpaceVLN introduces an efficient stagewise closed-loop framework where planning and execution are organized around verifiable space--landmark stages. During navigation, the agent progressively abstracts explored regions into Spatial Waypoints and dynamically maintains subtask-grounded landmark evidence, forming a hierarchical Spatial Cognitive Memory for progress localization and spatial-relation understanding. Built on this memory, Spatial-CoT integrates task-progress reasoning with spatial perception, analysis, and prediction, enabling Task-Guided Spatial Reasoning for embodied navigation. The unified stage interface enables SpaceVLN to address both Vision-and-Language Navigation and Object-Goal Navigation under a unified zero-shot setting, without task-specific policy training. Across R2R-CE, RxR-CE, GN-Bench, and HM3D-OVON, SpaceVLN achieves state-of-the-art zero-shot performance, and real-robot deployment further validates its applicability. These results highlight Spatial Cognitive Memory and Task-Guided Spatial Reasoning as a practical foundation for stronger embodied navigation agents.

具身智能导航空间推理零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。