arXiv:2603.09163cs.RO2026-03被引 5

让机器人看视频就能懂空间,导航更准更稳。

SPAN-Nav: Generalized Spatial Awareness for Versatile Vision-Language Navigation

  • 用视频预测空间占用,生成通用3D空间先验
  • 仅用一个令牌就捕捉关键空间线索,计算高效
  • 支持复杂场景真实部署,适合多任务导航应用

近期基于视觉语言模型的具身导航方法在多任务视觉语言导航中表现出强泛化能力,但在复杂环境中仍面临路径规划不可靠的问题,根源在于空间感知不足。本文提出SPAN-Nav,一种端到端的基础模型,通过处理RGB视频流,为具身导航注入通用3D空间意识。该模型在涵盖室内外多种场景的大规模数据上进行占用预测训练,提取跨场景的空间先验。为降低计算开销,研究发现仅需一个紧凑令牌即可捕获导航所需的粗粒度空间线索。受思维链(CoT)启发,该空间令牌被直接注入动作推理过程,实现端到端的空间引导。通过多任务协同训练,模型从通用空间先验中学习任务自适应线索,即使在缺乏显式空间监督的任务中也能保持鲁棒空间感知。为此,我们构建了包含420万条占用标注的超大规模数据集,覆盖多种导航任务类型。SPAN-Nav在三个跨场景、多任务基准测试中均达到当前最优性能,真实世界实验进一步验证了其在复杂物理环境中的泛化能力与实际可靠性。

原文摘要 · Abstract (English)

Recent embodied navigation approaches leveraging Vision-Language Models (VLMs) demonstrate strong generalization in versatile Vision-Language Navigation (VLN). However, reliable path planning in complex environments remains challenging due to insufficient spatial awareness. In this work, we introduce SPAN-Nav, an end-to-end foundation model designed to infuse embodied navigation with universal 3D spatial awareness using RGB video streams. SPAN-Nav extracts spatial priors across diverse scenes through an occupancy prediction task on extensive indoor and outdoor environments. To mitigate the computational burden, we introduce a compact representation for spatial priors, finding that a single token is sufficient to encapsulate the coarse-grained cues essential for navigation tasks. Furthermore, inspired by the Chain-of-Thought (CoT) mechanism, SPAN-Nav utilizes this single spatial token to explicitly inject spatial cues into action reasoning through an end-to end framework. Leveraging multi-task co-training, SPAN-Nav captures task-adaptive cues from generalized spatial priors, enabling robust spatial awareness to generalize even to the task lacking explicit spatial supervision. To support comprehensive spatial learning, we present a massive dataset of 4.2 million occupancy annotations that covers both indoor and outdoor scenes across multi-type navigation tasks. SPAN-Nav achieves state-of-the-art performance across three benchmarks spanning diverse scenarios and varied navigation tasks. Finally, real-world experiments validate the robust generalization and practical reliability of our approach across complex physical scenarios.

具身智能空间感知视觉导航多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。