P³Nav统一感知、预测与规划,提升视觉语言导航的场景理解能力。
P$^{3}$Nav: End-to-End Perception, Prediction and Planning for Vision-and-Language Navigation
- 融合物体级与地图级视觉线索增强感知
- 预测未来路径点,实现前瞻性决策
- 端到端设计适合复杂指令导航任务
在视觉-语言导航(VLN)中,智能体需根据语言指令,利用视觉观察规划通向目标的路径。现有方法主要通过视觉-文本对齐构建强大规划器,但常忽略规划前的全面场景理解,导致感知与预测能力不足。为此,我们提出P³Nav,一种整合感知、预测与规划的端到端框架,以增强智能体的场景理解并提升导航成功率。具体而言,P³Nav通过物体级与地图级视角提取互补线索,增强感知;进而预测未来路径点,使智能体具备对候选位置的内在认知;基于这些未来路径点,进一步预测语义地图信息,实现主动规划,减少对历史上下文的依赖。最后,综合感知与预测线索,由全局规划模块完成导航任务。大量实验表明,P³Nav在REVERIE、R2R-CE和RxR-CE基准上均达到新最优性能。
原文摘要 · Abstract (English)
In Vision-and-Language Navigation (VLN), an agent is required to plan a path to the target specified by the language instruction, using its visual observations. Consequently, prevailing VLN methods primarily focus on building powerful planners through visual-textual alignment. However, these approaches often bypass the imperative of comprehensive scene understanding prior to planning, leaving the agent with insufficient perception or prediction capabilities. Thus, we propose P$^{3}$Nav, a novel end-to-end framework integrating perception, prediction, and planning in a unified pipeline to strengthen the VLN agent's scene understanding and boost navigation success. Specifically, P$^{3}$Nav augments perception by extracting complementary cues from object-level and map-level perspectives. Subsequently, our P$^{3}$Nav predicts waypoints to model the agent's potential future states, endowing the agent with intrinsic awareness of candidate positions during navigation. Conditioned on these future waypoints, P$^{3}$Nav further forecasts semantic map cues, enabling proactive planning and reducing the strict reliance on purely historical context. Integrating these perceptual and predictive cues, a holistic planning module finally carries out the VLN tasks. Extensive experiments demonstrate that our P$^{3}$Nav achieves new state-of-the-art performance on the REVERIE, R2R-CE, and RxR-CE benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。