arXiv:2603.03739cs.CVcs.AI2026-03被引 2

通过融合语义与空间信息,实现更鲁棒的实时视觉语言导航

PROSPECT: Unified Streaming Vision-Language Navigation via Semantic--Spatial Fusion and Latent Predictive Representation

  • 用流式3D空间编码器和语义特征交叉注意力融合环境信息
  • 在冻结模型的潜在空间中预测下一步2D/3D特征,提升长程导航能力
  • 适用于真实机器人部署,对光照变化有更强适应性

多模态大语言模型推动了零样本端到端视觉语言导航的发展,但鲁棒导航不仅需要语义理解,还需对环境动态和空间结构进行预测建模。我们提出PROSPECT,一种统一的流式导航智能体,将流式视觉-语言-动作(VLA)策略与潜在预测表示学习相结合。PROSPECT采用CUT3R作为流式3D基础空间编码器,生成长上下文、绝对尺度的空间特征,并通过交叉注意力将其与SigLIP语义特征融合。训练时引入可学习的流查询令牌,从流式上下文中查询并预测下一步的2D和3D潜在特征(而非像素或显式模态),监督信号来自冻结的SigLIP和CUT3R教师模型的潜在空间。该预测分支在不增加推理开销的前提下塑造内部表示。在VLN-CE基准和真实机器人部署上的实验表明,PROSPECT达到最先进性能,在多种光照条件下展现出更优的长时程鲁棒性。代码即将开源。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have advanced zero-shot end-to-end Vision-Language Navigation (VLN), yet robust navigation requires not only semantic understanding but also predictive modeling of environment dynamics and spatial structure. We propose PROSPECT, a unified streaming navigation agent that couples a streaming Vision-Language-Action (VLA) policy with latent predictive representation learning. PROSPECT uses CUT3R as a streaming 3D foundation spatial encoder to produce long-context, absolute-scale spatial features, and fuses them with SigLIP semantic features via cross-attention. During training, we introduce learnable stream query tokens that query the streaming context and predict next-step 2D and 3D latent features (rather than pixels or explicit modalities), supervised in the latent spaces of frozen SigLIP and CUT3R teachers. The predictive branch shapes internal representations without inference overhead. Experiments on VLN-CE benchmarks and real-robot deployment demonstrate state-of-the-art performance and improved long-horizon robustness under diverse lighting. We will release code for the community soon.

视觉导航多模态流式处理机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。