通过构建空间场景图,实现零样本视觉语言导航的高效探索与定位。
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
- 构建空间场景图显式捕捉全局空间结构与语义信息。
- 在离散与连续环境中均显著优于现有零样本方法。
- 适合需要泛化能力的零样本导航任务研究者。
尽管基于学习的视觉-语言导航(VLN)代理可以从大规模训练数据中隐式学习空间知识,但零样本VLN代理缺乏这一过程,主要依赖局部观测进行导航,导致探索效率低且性能差距显著。为此,我们提出一种允许代理在任务执行前完全探索环境的零样本VLN设置,并构建空间场景图(SSG),以显式捕捉所探索环境中的全局空间结构与语义。基于此,我们提出SpatialNav,一个集成代理中心空间地图、罗盘对齐视觉表示和远程目标定位策略的零样本VLN代理,实现高效导航。在离散与连续环境中的综合实验表明,SpatialNav显著优于现有零样本代理,并明显缩小了与先进学习型方法的差距。这些结果凸显了全局空间表征对可泛化导航的重要性。
原文摘要 · Abstract (English)
Although learning-based vision-and-language navigation (VLN) agents can learn spatial knowledge implicitly from large-scale training data, zero-shot VLN agents lack this process, relying primarily on local observations for navigation, which leads to inefficient exploration and a significant performance gap. To deal with the problem, we consider a zero-shot VLN setting that agents are allowed to fully explore the environment before task execution. Then, we construct the Spatial Scene Graph (SSG) to explicitly capture global spatial structure and semantics in the explored environment. Based on the SSG, we introduce SpatialNav, a zero-shot VLN agent that integrates an agent-centric spatial map, a compass-aligned visual representation, and a remote object localization strategy for efficient navigation. Comprehensive experiments in both discrete and continuous environments demonstrate that SpatialNav significantly outperforms existing zero-shot agents and clearly narrows the gap with state-of-the-art learning-based methods. Such results highlight the importance of global spatial representations for generalizable navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。