arXiv:2601.18188cs.CVcs.AI2026-01被引 5

让AI导航更懂动作对视觉的影响,提升稳定性和规划能力。

\textsc{NaVIDA}: Vision-Language Navigation with Inverse Dynamics Augmentation

  • 用逆动力学监督显式建模动作如何改变视觉,增强策略学习
  • 在任务中实现比顶尖方法更好性能,参数量减少62.5%
  • 适合需要长程规划和鲁棒导航的智能体研究与应用

视觉-语言导航(VLN)要求智能体根据自然语言指令在视觉丰富的环境中作出连贯行动。然而,现有方法大多依赖反应式状态-动作映射,缺乏对动作如何影响后续视觉观察的显式建模。由于无法理解动作带来的视觉变化,智能体难以合理规划,导致行为不稳定、泛化能力弱及轨迹累积误差。为此,本文提出 extsc{NaVIDA}(Navigation with Inverse Dynamics Augmentation),一个轻量级的VLN框架,通过引入逆动力学监督(IDS)作为显式目标,将动作-视觉动态关系嵌入策略学习中。通过在共享表示和动作空间中联合优化视觉动态与指令条件的动作预测, extsc{NaVIDA} 提供结构化监督,正则化学习过程,提升导航稳定性与一致性。为强化监督并扩展有效规划范围, extsc{NaVIDA} 采用分层概率动作分块(HPAC),将轨迹组织为多步动作块,提供更具判别性的长程视觉变化提示。大量实验表明, extsc{NaVIDA} 在性能上超越现有最优方法,仅使用30亿参数(相比80亿参数),同时在真实机器人上验证了其实际可行性与有效性。

原文摘要 · Abstract (English)

Vision-and-Language Navigation (VLN) requires agents to interpret natural language instructions and act coherently in visually rich environments. However, most existing methods rely on reactive state-action mappings without explicitly action-grounded visual dynamics modeling. Lacking awareness of how actions transform subsequent visual observations, agents cannot plan actions rationally, leading to unstable behaviors, weak generalization, and cumulative error along trajectory. To address these issues, we introduce \textsc{NaVIDA} (\textbf{Nav}igation with \textbf{I}nverse \textbf{D}ynamics \textbf{A}ugmentation), a lightweight VLN framework that incorporates inverse dynamics supervision (IDS) as an explicit objective to embed action-grounded visual dynamics into policy learning. By jointly optimizing this visual dynamics with instruction-conditioned action prediction in a shared representation and action space, \textsc{NaVIDA} provides additional structured supervision that regularizes learning and leads to more stable and consistent navigation. To structure this supervision and extend the effective planning range, \textsc{NaVIDA} employs hierarchical probabilistic action chunking (HPAC), which organizes trajectories into multi-step chunks and provides discriminative, longer-range visual-change cues. Extensive experiments show that \textsc{NaVIDA} achieves superior navigation performance compared to state-of-the-art methods with fewer parameters (3B vs. 8B). Real-world robot evaluations further validate the practical feasibility and effectiveness of our approach.

视觉导航逆动力学动作规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。