arXiv:2412.08591cs.CVcs.AI2024-12CVPR被引 27

用真实房间视频构建3D导航数据集,提升智能体开放世界导航能力

RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation

  • 基于网络房间视频生成带3D重建的步行轨迹与自然语言指令
  • 包含约10万条开放指令轨迹,支持多任务性能显著提升
  • 适合研究开放世界导航与零样本训练的学者使用

视觉-语言导航(VLN)受限于训练数据的规模与多样性,主要源于现有模拟器需人工标注。为此,我们提出RoomTour3D,一个源自网络房间巡展视频的视频-指令数据集,真实捕捉室内空间与人类行走示范。不同于已有数据集,其利用在线视频的规模与多样性,生成开放式人行轨迹与开放世界导航指令。为弥补在线视频中缺乏导航数据的问题,我们进行3D重建,获得包含房间类型、物体位置及周围场景三维形状信息的3D行走路径。数据集包含约10万条描述丰富的开放轨迹和约20万条指令,以及来自1847个房间环境的1.7万条动作丰富轨迹。实验表明,RoomTour3D在多个VLN任务(如CVDN、SOON、R2R、REVERIE)中均实现显著性能提升,并推动可训练的零样本导航智能体发展,揭示了迈向开放世界导航的潜力与挑战。

原文摘要 · Abstract (English)

Vision-and-Language Navigation (VLN) suffers from the limited diversity and scale of training data, primarily constrained by the manual curation of existing simulators. To address this, we introduce RoomTour3D, a video-instruction dataset derived from web-based room tour videos that capture real-world indoor spaces and human walking demonstrations. Unlike existing VLN datasets, RoomTour3D leverages the scale and diversity of online videos to generate open-ended human walking trajectories and open-world navigable instructions. To compensate for the lack of navigation data in online videos, we perform 3D reconstruction and obtain 3D trajectories of walking paths augmented with additional information on the room types, object locations and 3D shape of surrounding scenes. Our dataset includes $\sim$100K open-ended description-enriched trajectories with $\sim$200K instructions, and 17K action-enriched trajectories from 1847 room tour environments. We demonstrate experimentally that RoomTour3D enables significant improvements across multiple VLN tasks including CVDN, SOON, R2R, and REVERIE. Moreover, RoomTour3D facilitates the development of trainable zero-shot VLN agents, showcasing the potential and challenges of advancing towards open-world navigation.

视觉导航3D重建指令学习开放世界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。