无需全景图和分步规划,用三张正视图实现快速零样本导航
Fast-SmartWay: Panoramic-Free End-to-End Zero-Shot Vision-and-Language Navigation
- 仅用三张正向RGB-D图与语言指令,端到端直接预测动作
- 每步延迟显著降低,实测表现优于或媲美全景基线
- 引入不确定性感知推理,提升决策鲁棒性与全局规划能力
连续环境中的视觉语言导航(VLN-CE)近年借助多模态大语言模型(MLLMs)实现了零样本导航。然而现有方法通常依赖全景观测和包含航点预测的两阶段流程,导致显著延迟,限制了实际应用。本文提出Fast-SmartWay,一种端到端零样本VLN-CE框架,无需全景视图与航点预测器。该方法仅使用三张前向RGB-D图像与自然语言指令,使MLLM直接生成动作。为增强决策鲁棒性,引入不确定性感知推理模块,包含:(i) 消歧模块以避免局部最优,(ii) 前后双向推理机制实现全局一致规划。在模拟与真实机器人环境上的实验表明,本方法显著降低每步延迟,同时性能达到或超越全景视图基线。结果证明Fast-SmartWay在真实世界零样本具身导航中的实用性与有效性。
原文摘要 · Abstract (English)
Recent advances in Vision-and-Language Navigation in Continuous Environments (VLN-CE) have leveraged multimodal large language models (MLLMs) to achieve zero-shot navigation. However, existing methods often rely on panoramic observations and two-stage pipelines involving waypoint predictors, which introduce significant latency and limit real-world applicability. In this work, we propose Fast-SmartWay, an end-to-end zero-shot VLN-CE framework that eliminates the need for panoramic views and waypoint predictors. Our approach uses only three frontal RGB-D images combined with natural language instructions, enabling MLLMs to directly predict actions. To enhance decision robustness, we introduce an Uncertainty-Aware Reasoning module that integrates (i) a Disambiguation Module for avoiding local optima, and (ii) a Future-Past Bidirectional Reasoning mechanism for globally coherent planning. Experiments on both simulated and real-robot environments demonstrate that our method significantly reduces per-step latency while achieving competitive or superior performance compared to panoramic-view baselines. These results demonstrate the practicality and effectiveness of Fast-SmartWay for real-world zero-shot embodied navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。