构建首个视觉语言导航概念标注数据集,支持智能体路径理解研究。
NAVCON: A Cognitively Inspired and Linguistically Grounded Corpus for Vision and Language Navigation
- 基于认知与语言学设计四类导航核心概念,自动生成大规模银标数据
- 包含约30万条指令的23.6万次概念标注和270万张对齐视觉图像
- 适用于导航模型训练与大模型少样本学习,推动智能体理解自然指令
我们提出NAVCON,一个基于R2R和RxR两个主流数据集构建的大规模视觉-语言导航(VLN)语料库。论文定义了四个受认知启发且语言基础牢固的导航概念,并设计算法生成这些概念在自然导航指令中的大规模银标标注。每条标注指令均配以智能体执行时的视频片段,展示其视觉感知。NAVCON包含约30万条指令的23.6万次概念标注,以及约1.9万条指令对应的270万张对齐图像。我们通过人工评估验证了银标质量,还训练了一个检测模型识别未见指令中的导航概念及其语言表达。此外,使用GPT-4o进行少样本学习,在该数据集上表现良好。
原文摘要 · Abstract (English)
We present NAVCON, a large-scale annotated Vision-Language Navigation (VLN) corpus built on top of two popular datasets (R2R and RxR). The paper introduces four core, cognitively motivated and linguistically grounded, navigation concepts and an algorithm for generating large-scale silver annotations of naturally occurring linguistic realizations of these concepts in navigation instructions. We pair the annotated instructions with video clips of an agent acting on these instructions. NAVCON contains 236, 316 concept annotations for approximately 30, 0000 instructions and 2.7 million aligned images (from approximately 19, 000 instructions) showing what the agent sees when executing an instruction. To our knowledge, this is the first comprehensive resource of navigation concepts. We evaluated the quality of the silver annotations by conducting human evaluation studies on NAVCON samples. As further validation of the quality and usefulness of the resource, we trained a model for detecting navigation concepts and their linguistic realizations in unseen instructions. Additionally, we show that few-shot learning with GPT-4o performs well on this task using large-scale silver annotations of NAVCON.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。