arXiv:2409.02389cs.CVcs.AI2024-09NeurIPS被引 64

构建大规模多模态3D场景推理数据集,提升智能体对复杂环境的理解能力。

Multi-modal Situated Reasoning in 3D Scenes

  • 利用3D场景图与视觉语言模型,自动构建多模态问答数据
  • 包含25.1万组问题答案对,覆盖9类复杂场景任务
  • 支持文本、图像、点云多模态输入,适合训练具身智能模型

在具身智能中,情境感知对于理解与推理3D场景至关重要。然而,现有数据集和基准在模态多样性、数据规模与任务范围上均存在局限。为此,我们提出多模态情境问答(MSQA)数据集,通过3D场景图与视觉语言模型,在多样化真实3D场景中规模化构建该数据集。MSQA包含25.1万组情境问答对,涵盖9种不同问题类别,覆盖3D场景中的复杂情形。我们在基准中引入新颖的多模态交错输入方式,同时提供文本、图像与点云信息以描述情境与问题,解决了以往单模态(如仅文本)带来的歧义问题。此外,我们设计了多模态情境下一步导航(MSNN)基准,用于评估模型的导航情境推理能力。在MSQA与MSNN上的全面评估揭示了现有视觉语言模型的局限性,强调了处理多模态交错输入与情境建模的重要性。数据扩展与跨域迁移实验进一步证明,使用MSQA作为预训练数据集可有效提升情境推理模型性能。

原文摘要 · Abstract (English)

Situation awareness is essential for understanding and reasoning about 3D scenes in embodied AI agents. However, existing datasets and benchmarks for situated understanding are limited in data modality, diversity, scale, and task scope. To address these limitations, we propose Multi-modal Situated Question Answering (MSQA), a large-scale multi-modal situated reasoning dataset, scalably collected leveraging 3D scene graphs and vision-language models (VLMs) across a diverse range of real-world 3D scenes. MSQA includes 251K situated question-answering pairs across 9 distinct question categories, covering complex scenarios within 3D scenes. We introduce a novel interleaved multi-modal input setting in our benchmark to provide text, image, and point cloud for situation and question description, resolving ambiguity in previous single-modality convention (e.g., text). Additionally, we devise the Multi-modal Situated Next-step Navigation (MSNN) benchmark to evaluate models' situated reasoning for navigation. Comprehensive evaluations on MSQA and MSNN highlight the limitations of existing vision-language models and underscore the importance of handling multi-modal interleaved inputs and situation modeling. Experiments on data scaling and cross-domain transfer further demonstrate the efficacy of leveraging MSQA as a pre-training dataset for developing more powerful situated reasoning models.

3D场景多模态具身智能问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。