提升视觉语言导航的零样本表现,增强路径预测与回溯能力
SmartWay: Enhanced Waypoint Prediction and Backtracking for Zero-Shot Vision-and-Language Navigation
- 用更强视觉编码器和占用感知损失优化路径点生成
- 结合多模态大模型实现带历史记忆的自适应路径规划
- 在真实机器人上验证,适合零样本场景下的复杂导航
连续环境中的视觉-语言导航(VLN)要求智能体在未受约束的3D空间中理解自然语言指令并导航。现有VLN-CE框架采用两阶段方法:先生成路径点,再执行移动。但当前路径点预测器缺乏空间感知,导航器则缺少历史推理与回溯能力,限制了适应性。本文提出一种零样本VLN-CE框架,集成增强型路径点预测器与基于多模态大语言模型(MLLM)的导航器。预测器采用更强的视觉编码器、掩码交叉注意力融合及占用感知损失,提升路径点质量;导航器引入历史感知推理与自适应路径规划,支持回溯,增强鲁棒性。在R2R-CE与MP3D基准测试中,本方法在零样本设置下达到当前最优性能,表现媲美全监督方法。在Turtlebot 4上的真实世界验证进一步证明其适应性。
原文摘要 · Abstract (English)
Vision-and-Language Navigation (VLN) in continuous environments requires agents to interpret natural language instructions while navigating unconstrained 3D spaces. Existing VLN-CE frameworks rely on a two-stage approach: a waypoint predictor to generate waypoints and a navigator to execute movements. However, current waypoint predictors struggle with spatial awareness, while navigators lack historical reasoning and backtracking capabilities, limiting adaptability. We propose a zero-shot VLN-CE framework integrating an enhanced waypoint predictor with a Multi-modal Large Language Model (MLLM)-based navigator. Our predictor employs a stronger vision encoder, masked cross-attention fusion, and an occupancy-aware loss for better waypoint quality. The navigator incorporates history-aware reasoning and adaptive path planning with backtracking, improving robustness. Experiments on R2R-CE and MP3D benchmarks show our method achieves state-of-the-art (SOTA) performance in zero-shot settings, demonstrating competitive results compared to fully supervised methods. Real-world validation on Turtlebot 4 further highlights its adaptability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。