构建了超大规模3D场景语义理解与导航数据集,支持自然语言指令下的智能导航。
VLA-3D: A Dataset for 3D Semantic Scene Understanding and Navigation
- 整合11.5K真实3D室内场景,生成2350万条物体语义关系与970万条指代描述。
- 包含点云、语义标注、场景图和可导航空间信息,支持视点无关的空间关系理解。
- 面向开放世界变化场景,适合研究鲁棒性导航与多模态交互系统的研究者使用。
随着大语言模型(LLMs)、视觉-语言模型(VLMs)等通用基础模型的发展,基于自然语言输入在多样化环境中运行的多模态、多任务具身智能体展现出巨大潜力。其中,基于自然语言指令的室内导航是一个关键应用方向。然而,由于需要复杂的空间推理与细粒度语义理解,尤其在包含大量细分类别物体的任意场景中,该问题仍具挑战性。为此,我们构建了目前最大的真实世界3D场景视觉-语言引导动作数据集(VLA-3D),涵盖超过11.5K个从现有数据集中扫描的3D室内房间,生成2350万条物体间语义关系,以及970万条合成的指代语句。数据集包含处理后的3D点云、语义物体与房间标注、场景图、可导航自由空间标注,以及聚焦于视点无关空间关系的指代语言表述,旨在提升下游导航任务的性能,尤其在面对开放世界中场景变化与语言不完美时具备更强鲁棒性。我们使用当前最先进的模型对该数据集进行基准测试,建立性能基线。所有数据生成与可视化代码已公开,详见 https://github.com/HaochenZ11/VLA-3D。随着该数据集发布,我们希望为鲁棒的3D语义场景理解提供资源,并推动交互式室内导航系统的研发。
原文摘要 · Abstract (English)
With the recent rise of Large Language Models (LLMs), Vision-Language Models (VLMs), and other general foundation models, there is growing potential for multimodal, multi-task embodied agents that can operate in diverse environments given only natural language as input. One such application area is indoor navigation using natural language instructions. However, despite recent progress, this problem remains challenging due to the spatial reasoning and semantic understanding required, particularly in arbitrary scenes that may contain many objects belonging to fine-grained classes. To address this challenge, we curate the largest real-world dataset for Vision and Language-guided Action in 3D Scenes (VLA-3D), consisting of over 11.5K scanned 3D indoor rooms from existing datasets, 23.5M heuristically generated semantic relations between objects, and 9.7M synthetically generated referential statements. Our dataset consists of processed 3D point clouds, semantic object and room annotations, scene graphs, navigable free space annotations, and referential language statements that specifically focus on view-independent spatial relations for disambiguating objects. The goal of these features is to aid the downstream task of navigation, especially on real-world systems where some level of robustness must be guaranteed in an open world of changing scenes and imperfect language. We benchmark our dataset with current state-of-the-art models to obtain a performance baseline. All code to generate and visualize the dataset is publicly released, see https://github.com/HaochenZ11/VLA-3D. With the release of this dataset, we hope to provide a resource for progress in semantic 3D scene understanding that is robust to changes and one which will aid the development of interactive indoor navigation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。