用轨迹规划与主动想象实现零样本视觉语言导航,更智能、更省资源。
DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation
- 通过视角校正降低感知成本,稳定第一人称视觉输入。
- 采用轨迹级规划,使动作语义更契合指令意图。
- 引入想象预测器,支持长远预判,适合复杂场景导航。
视觉-语言导航在连续环境(VLN-CE)中将语言指令与感知控制相连接,是具身机器人的核心能力。近期大模型被用作感知、推理与行动的通用先验,实现了无需任务微调的零样本导航。然而,现有方法依赖高成本感知与被动场景理解,仅能进行点级动作决策,导致部署昂贵、动作语义错位且规划短视。为此,我们提出DreamNav:(1)EgoView Corrector对齐视角,稳定第一人称感知;(2)轨迹预测器支持全局轨迹规划,更好匹配指令语义;(3)想象预测器赋予代理主动思考能力,实现前瞻式长程规划。在VLN-CE与真实世界测试中,DreamNav达到新的零样本最佳性能,相比最强基线(含额外信息)在成功率(SR)和路径相似率(SPL)上分别提升7.49%和18.15%。据我们所知,这是首个仅使用第一人称输入却统一轨迹规划与主动想象的零样本导航方法。
原文摘要 · Abstract (English)
Vision-and-Language Navigation in Continuous Environments (VLN-CE), which links language instructions to perception and control in the real world, is a core capability of embodied robots. Recently, large-scale pretrained foundation models have been leveraged as shared priors for perception, reasoning, and action, enabling zero-shot VLN without task-specific training. However, existing zero-shot VLN methods depend on costly perception and passive scene understanding, collapsing control to point-level choices. As a result, they are expensive to deploy, misaligned in action semantics, and short-sighted in planning. To address these issues, we present DreamNav that focuses on the following three aspects: (1) for reducing sensory cost, our EgoView Corrector aligns viewpoints and stabilizes egocentric perception; (2) instead of point-level actions, our Trajectory Predictor favors global trajectory-level planning to better align with instruction semantics; and (3) to enable anticipatory and long-horizon planning, we propose an Imagination Predictor to endow the agent with proactive thinking capability. On VLN-CE and real-world tests, DreamNav sets a new zero-shot state-of-the-art (SOTA), outperforming the strongest egocentric baseline with extra information by up to 7.49\% and 18.15\% in terms of SR and SPL metrics. To our knowledge, this is the first zero-shot VLN method to unify trajectory-level planning and active imagination while using only egocentric inputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。