arXiv:2503.17406cs.CVcs.RO2025-03ICRA被引 8

构建首个面向3D场景中不完美语言的交互式指代定位基准数据集。

IRef-VLA: A Benchmark for Interactive Referential Grounding with Imperfect Language in 3D Scenes

  • 构建包含11.5万+3D房间的真实世界数据集,含470万条不完美语言描述
  • 首次在真实3D场景中系统评估语言不准确时的指代定位性能
  • 适合研究多模态机器人导航、鲁棒自然语言理解的研究者

随着大语言模型和视觉-语言模型的发展,基于自然语言输入在多样化环境中运行的多模态、多任务机器人具备巨大潜力。其中一个关键应用是室内导航。然而,由于需要3D空间推理与语义理解,该任务仍具挑战性,且语言可能不准确或与场景不符,进一步增加难度。为此,我们构建了IRef-VLA基准数据集,用于3D场景中的交互式指代视觉-语言引导行动,支持不完美语言输入。IRef-VLA是目前最大的真实世界指代定位数据集,包含超过11.5K个扫描3D房间,760万条启发式生成的语义关系,以及470万条指代陈述。数据集还包含语义物体与房间标注、场景图、可通行空间标注,并增强包含语言缺陷或歧义的语句。我们通过评估先进模型建立性能基线,并开发图搜索基线以展示性能上限及利用场景图生成替代方案。该基准旨在推动鲁棒、交互式导航系统的研发。数据集与全部源码已公开于https://github.com/HaochenZ11/IRef-VLA。

原文摘要 · Abstract (English)

With the recent rise of large language models, vision-language models, and other general foundation models, there is growing potential for multimodal, multi-task robotics that can operate in diverse environments given natural language input. One such application is indoor navigation using natural language instructions. However, despite recent progress, this problem remains challenging due to the 3D spatial reasoning and semantic understanding required. Additionally, the language used may be imperfect or misaligned with the scene, further complicating the task. To address this challenge, we curate a benchmark dataset, IRef-VLA, for Interactive Referential Vision and Language-guided Action in 3D Scenes with imperfect references. IRef-VLA is the largest real-world dataset for the referential grounding task, consisting of over 11.5K scanned 3D rooms from existing datasets, 7.6M heuristically generated semantic relations, and 4.7M referential statements. Our dataset also contains semantic object and room annotations, scene graphs, navigable free space annotations, and is augmented with statements where the language has imperfections or ambiguities. We verify the generalizability of our dataset by evaluating with state-of-the-art models to obtain a performance baseline and also develop a graph-search baseline to demonstrate the performance bound and generation of alternatives using scene-graph knowledge. With this benchmark, we aim to provide a resource for 3D scene understanding that aids the development of robust, interactive navigation systems. The dataset and all source code is publicly released at https://github.com/HaochenZ11/IRef-VLA.

3D导航指代定位多模态机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。