arXiv:2411.11394cs.RO2024-11被引 9

用YouTube视频自动生成高质量视觉语言导航指令。

InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models

  • 基于大模型从视频中自动提取路径与指令对
  • 在R2R/RxR上实现领先性能,尤其在未知环境
  • 适合研究视觉语言导航与数据生成的学者

视觉-语言导航(VLN)中的智能体因缺乏真实训练环境和高质量路径-指令对,导致在未知环境中泛化能力差。现有构建真实场景的方法成本高,指令扩展依赖预设模板,适应性弱。为此,我们提出InstruGen,一种基于大视觉语言模型(LMMs)的自动路径-指令对生成范式。我们采用YouTube房屋巡视频道作为真实导航场景,利用LMM强大的视觉理解与生成能力,自动生成多样且高质量的路径-指令对。该方法能生成不同粒度的导航指令,并实现指令与视觉观测的细粒度对齐,这是以往方法难以做到的。此外,我们设计多阶段验证机制以降低LMM的幻觉和不一致性。实验表明,使用InstruGen生成的数据训练的智能体在R2R和RxR基准上达到当前最优表现,尤其在未见环境中表现突出。代码已开源:https://github.com/yanyu0526/InstruGen。

原文摘要 · Abstract (English)

Recent research on Vision-and-Language Navigation (VLN) indicates that agents suffer from poor generalization in unseen environments due to the lack of realistic training environments and high-quality path-instruction pairs. Most existing methods for constructing realistic navigation scenes have high costs, and the extension of instructions mainly relies on predefined templates or rules, lacking adaptability. To alleviate the issue, we propose InstruGen, a VLN path-instruction pairs generation paradigm. Specifically, we use YouTube house tour videos as realistic navigation scenes and leverage the powerful visual understanding and generation abilities of large multimodal models (LMMs) to automatically generate diverse and high-quality VLN path-instruction pairs. Our method generates navigation instructions with different granularities and achieves fine-grained alignment between instructions and visual observations, which was difficult to achieve with previous methods. Additionally, we design a multi-stage verification mechanism to reduce hallucinations and inconsistency of LMMs. Experimental results demonstrate that agents trained with path-instruction pairs generated by InstruGen achieves state-of-the-art performance on the R2R and RxR benchmarks, particularly in unseen environments. Code is available at https://github.com/yanyu0526/InstruGen.

视觉语言导航指令生成大模型应用数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。