arXiv:2505.03174cs.ROcs.CV2025-05被引 1

用手机GPS语音和NLP自动生成自动驾驶视觉语言导航数据。

Automated Data Curation Using GPS & NLP to Generate Instruction-Action Pairs for Autonomous Vehicle Vision-Language Navigation Datasets

  • 通过手机GPS语音与NLP自动提取指令-动作对。
  • 构建了8类指令分类,支持多样化的导航语义。
  • 适合研究自动驾驶视觉语言导航的数据集构建者。

指令-动作(IA)数据对对训练机器人系统,尤其是自动驾驶车辆(AV)具有重要价值,但人工标注成本高且效率低。本文探索利用移动应用中的全球定位系统(GPS)参考和自然语言处理(NLP)技术,无需人工生成或事后标记,即可自动构建大量IA指令与响应。在初步数据采集中,通过前往不同目的地并收集GPS应用的语音指令,我们展示了获取和分类多样化指令的方法,并结合视频数据形成完整的视觉-语言-动作三元组。本文详细介绍了全自动数据采集原型系统ADVLAT-Engine。我们将收集到的GPS语音指令划分为八类,凸显了来自免费移动应用的丰富命令与指代性内容。通过研究利用GPS参考实现IA数据对自动化的潜力,可显著提升高质量IA数据集的生成速度与规模,同时降低费用,为视觉-语言-导航(VLN)任务及人机交互式自主系统提供更鲁棒的视觉-语言-动作(VLA)模型基础。

原文摘要 · Abstract (English)

Instruction-Action (IA) data pairs are valuable for training robotic systems, especially autonomous vehicles (AVs), but having humans manually annotate this data is costly and time-inefficient. This paper explores the potential of using mobile application Global Positioning System (GPS) references and Natural Language Processing (NLP) to automatically generate large volumes of IA commands and responses without having a human generate or retroactively tag the data. In our pilot data collection, by driving to various destinations and collecting voice instructions from GPS applications, we demonstrate a means to collect and categorize the diverse sets of instructions, further accompanied by video data to form complete vision-language-action triads. We provide details on our completely automated data collection prototype system, ADVLAT-Engine. We characterize collected GPS voice instructions into eight different classifications, highlighting the breadth of commands and referentialities available for curation from freely available mobile applications. Through research and exploration into the automation of IA data pairs using GPS references, the potential to increase the speed and volume at which high-quality IA datasets are created, while minimizing cost, can pave the way for robust vision-language-action (VLA) models to serve tasks in vision-language navigation (VLN) and human-interactive autonomous systems.

自动驾驶视觉语言导航数据生成NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。