用视频先验生成精准导航指令,支持复杂路径规划。
Goal-oriented Navigation Instruction Generation with Tour Video Priors

- 基于视角视频与目标生成指令,无需地图等中间表示。
- 在6万段视频和3.7万个提示上验证,模型指令可执行性显著提升。
- 适合研究视觉语言导航、指令生成的学者使用。
导航指令生成(NIG)旨在为导航提供逐步的自然语言指导。现有研究多将NIG作为视觉-语言导航(VLN)的辅助任务,侧重数据增强或多任务学习。然而,仅从紧凑环境先验生成导航指令需精细的空间推理,尤其当目标路径不沿示范路线时,当前多模态模型仍面临挑战。本文提出VideoNIG:一种以目标为导向的视频引导式导航指令生成任务,基于第一视角游览视频、初始观察和文本或视觉目标生成指令,无需依赖图谱、地图等中间表示。我们在一个可控模拟器基准中构建了60,000段连续室内环境的游览视频和37,000个具有渐进难度的多模态提示。引入诊断评估协议,结合文本相似性、选择式空间一致性测试及下游导航执行能力。为此,提出两阶段课程学习框架,分解为基础运动感知与长程导航推理。首先通过动作预热实现空间动作-视角对齐,再通过轨迹复杂度递增进行探索难度推进。大量实验表明,现有多模态大模型在VideoNIG上表现不佳,而我们的方法在互补诊断指标上显著提升指令质量。最终,将生成指令集成至VLN代理,验证了该任务形式在端到端导航中的可执行性。
原文摘要 · Abstract (English)
Navigation Instruction Generation (NIG) aims to produce step-by-step natural language instructions for navigation guidance. Existing studies primarily treat NIG as an auxiliary task for vision-andlanguage navigation (VLN), focusing on data augmentation or multi-task learning. However, generating navigation instructions from compact environmental priors requires meticulous spatial reasoning, especially when the target route does not simply follow the demonstrated tour, and remains challenging for current multimodal models. In this work, we introduce VideoNIG, a goal-oriented video-grounded NIG task that generates navigation instructions from ego-centric tour videos, an initial observation, and a textual or visual goal, without relying on intermediate representations such as graphs and maps. We instantiate VideoNIG in a controlled simulator benchmark with 60K tour videos across continuous indoor environments and 37K multimodal prompts with progressive difficulty levels. We further introduce a diagnostic evaluation protocol that combines text similarity, choice-based spatial consistency tests, and downstream navigation execution. To address this task, we propose a two-stage Curriculum Learning framework that decomposes the learning into foundational motion perception and long-horizon navigation reasoning. Specifically, we first employ Action Warmup for spatial action-view alignment, followed by Complexity Progression using trajectories with increasing exploratory difficulty. Extensive experiments show that existing MLLMs struggle with VideoNIG, while our approach significantly improves instruction quality across complementary diagnostic metrics. Finally, integrating VideoNIG-generated instructions with a VLN agent demonstrates the executability of this task formulation for end-to-end navigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。