将连续环境中的视觉轨迹转为可导航提示,生成更准确的导航指令。
VTInstructor: Visual Trajectory Prompting for Navigation Instruction Generation in Continuous Environments

- 通过关键帧提取与视觉线索叠加,将模糊轨迹转为显式导航提示。
- 在R2R-CE和RxR-CE上超越基线0.357和0.109的CIDEr得分。
- 生成指令使机器人成功率提升至63.3%,适合人机交互与数据增强场景。
在连续环境中从第一视角RGB视频生成导航指令是人机交互与大规模数据集构建的重要挑战。以往方法依赖离散视角图与全景观测,轨迹结构明确;但在连续环境中,智能体仅接收密集RGB流,导致轨迹线索难以恢复。本文提出首个面向连续环境的视觉轨迹提示框架VTInstructor。核心思想是将隐式轨迹几何转化为显式视觉轨迹提示:EDTC将长视频序列压缩为导航关键帧,VTP在这些锚点上叠加路径、转向与目标线索,VTMod将轨迹信号注入视觉编码器,VT-GRPO在训练中进一步校准空间注入,全程无需导航图、预建地图或场景重建。在具有挑战性的R2R-CE和RxR-CE验证集未见场景上,VTInstructor在所有标准自然语言生成指标上达到新最优,相较最强基线分别提升+0.357和+0.109 CIDEr。除自动指标外,其生成指令使冻结跟随者成功率提升至63.3%,较最佳对比源高14.7个百分点,并在下游导航任务中带来+3%的成功率增益。
原文摘要 · Abstract (English)
Navigation instruction generation from ego-centric RGB video in continuous environments is an important yet challenging task for human-robot interaction and scalable dataset construction. Prior instruction generators assume discrete viewpoint graphs with panoramic observations, where trajectory structure is explicit; in continuous environments, however, the agent receives only a dense RGB stream, making trajectory cues difficult to recover. We propose VTInstructor, the first VLN instruction generation framework for continuous environments. Our key idea is to convert implicit trajectory geometry into explicit visual trajectory prompts: EDTC condenses long RGB trajectories into navigation-critical keyframes, VTP overlays path, turn, and goal cues onto these anchors, VTMod injects the resulting trajectory signals into the visual encoder, and VT-GRPO further calibrates this spatial injection during training, all without requiring a navigation graph, pre-built map, or scene reconstruction. On the challenging R2R-CE and RxR-CE Val Unseen benchmarks, VTInstructor sets a new state of the art across all standard NLG metrics, surpassing the strongest baseline by +0.357 CIDEr and +0.109 CIDEr, respectively. Beyond automatic metrics, VTInstructor-generated instructions raise a frozen follower's success rate to 63.3%, a +14.7 percentage-point gain over the best competing instruction source, and provide consistent data augmentation gains of +3 SR points on downstream navigation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。