用可执行轨迹替代孤立路标,提升视觉语言导航的可达性与一致性
Beyond Waypoints: A Trajectory-Centric Waypointing Paradigm for Vision-Language Navigation

- 将路标建模为可执行轨迹,避免不可达问题
- 在VLN-CE基准上超越基线模型,成功率达78.3%
- 适合需高精度路径规划的机器人导航任务
视觉语言导航在连续环境(VLN-CE)中要求智能体根据自然语言指令在类真实世界环境中导航。现有方法通常采用三阶段框架:路标预测器生成可导航路标,导航器选择最佳路标,低层控制器执行移动。然而,这种解耦范式常导致路标不可达或规划与执行不一致。本文提出一种新范式——轨迹路标(Trajectory Waypoint),将每个候选路标与可执行轨迹关联。为此,设计基于TSDF引导的扩散策略作为轨迹路标预测器,有效避障并确保路标可达。进一步提出轨迹增强型导航器,将关联轨迹作为额外信息注入规划过程,实现高层语义决策与底层执行的严格一致。在VLN-CE基准上的大量实验表明,该范式显著优于基线方法,成功率提升至78.3%。
原文摘要 · Abstract (English)
Vision-Language Navigation in Continuous Environments (VLN-CE) requires agents to follow natural-language instructions while navigating in real-world-like environments. Most VLN-CE approach\-es adopt a three-stage framework: a waypoint predictor proposes navigable waypoints, and a navigator selects the best waypoint, with a low-level controller executing the movement to it. However, this decoupled paradigm often leads to unreachable waypoints or inconsistencies between planning and control. In this work, instead of predicting isolated waypoints, we introduce a novel paradigm called Trajectory Waypoint, which grounds each candidate waypoint in an executable trajectory. To realize this, we design a Trajectory Waypoint Predictor formulated as a TSDF-guided diffusion policy, which steers trajectory generation away from obstacles, inherently ensuring the reachability of the predicted waypoints. We further propose a trajectory-enhanced navigator that injects the associated trajectory as additional information for planning, enabling strict consistency between high-level semantic decisions and low-level execution. Extensive experiments on the VLN-CE benchmark show that our Trajectory Waypoint paradigm achieves superior performance over the baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。