arXiv:2606.30367cs.RO2026-06被引 1

让AI导航更懂环境变化,一次建模世界与动作。

FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation

论文配图:FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation
图 1 · 摘自论文原文
  • 联合编码文本、视觉与空间信息,统一建模世界状态与动作
  • 在多个基准上以40亿参数达到顶尖性能,显著超越此前方法
  • 适合研究智能体长期决策与环境预测的学者使用

连续环境中的视觉-语言导航(VLN)要求智能体在第一人称观测中理解指令,并在长序列动作中保持空间认知。现有导航基础模型虽通过扩大视觉-语言模型取得进展,但多仅将导航视为直接动作生成,缺乏对世界状态的显式建模或未来演化预测。本文提出FutureNav,一种基于视觉-语言模型的统一世界-动作建模框架。该框架联合编码文本、视觉与空间特征,输入大语言模型,并优化四项目标:动作策略目标用于导航动作预测,逆向与正向动力学目标用于建模状态转移,未来生成目标用于预测未来空间状态。该统一架构在增强动作预测的同时显式建模世界,且不牺牲推理速度。大量实验表明,仅使用40亿参数的骨干网络,FutureNav在多个VLN基准上达到当前最优表现,显著优于先前方法,为未来世界-动作模型在VLN中的发展铺平道路。代码与模型将开源以支持后续研究。

原文摘要 · Abstract (English)

Vision-and-language navigation (VLN) in continuous environments requires an agent to ground instructions in egocentric observations while maintaining spatial understanding across long action sequences. Recent navigation foundation models have shown strong progress by scaling vision-language models, but they often learn navigation primarily as direct action generation, without explicitly modeling world states or predicting their future evolution. We introduce FutureNav, a VLM-based unified world-action modeling framework for vision-and-language navigation. Specifically, FutureNav jointly encodes text, visual, and spatial features and feeds them into the LLM, and optimizes four objectives for simultaneous world and action modeling: an action policy objective for navigation action prediction, inverse and forward dynamics objectives for modeling state transitions, and a future generation objective for predicting future spatial states. This unified architecture strengthens action prediction while explicitly modeling the world, without sacrificing inference speed. Extensive experiments show that, with only a 4B-scale backbone, FutureNav achieves state-of-the-art performance on multiple VLN benchmarks and substantially outperforms prior VLN methods, paving the way toward future world-action models for VLN. We will release the code and models to support future research.

视觉导航世界建模语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。