arXiv:2507.12026cs.CV2025-07中稿 · IROS 2025被引 13

用大模型生成6万多组3D场景问答数据,提升智能体问答能力。

3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering

  • 融合多模态嵌入与跨模态交互,统一处理语言与3D场景
  • 在1513个场景中生成6.2万对问答和7.3万条物体描述
  • 在ScanQA/ScanRefer上分别提升CIDEr 2.15%和1.84%,适合场景理解研究

随着室内场景任务对多样化、可扩展数据的需求增加,我们提出3D-MoRe,一种利用基础模型优势生成大规模3D-语言数据集的新范式。该框架整合多模态嵌入、跨模态交互与语言模型解码器,处理自然语言指令与3D场景数据,增强复杂3D环境中的推理与生成能力。基于ScanNet 3D场景数据集,结合ScanQA与ScanRefer的文本标注,3D-MoRe在1,513个场景中生成62,000个问答对与73,000条物体描述。通过多种数据增强与语义过滤技术确保数据质量。在ScanQA上的实验表明,3D-MoRe显著优于现有基线,CIDEr得分提升2.15%;在ScanRefer上,[email protected]提升1.84%,验证了其在两项任务中的有效性。代码与生成数据集将公开发布于https://3D-MoRe.github.io。

原文摘要 · Abstract (English)

With the growing need for diverse and scalable data in indoor scene tasks, such as question answering and dense captioning, we propose 3D-MoRe, a novel paradigm designed to generate large-scale 3D-language datasets by leveraging the strengths of foundational models. The framework integrates key components, including multi-modal embedding, cross-modal interaction, and a language model decoder, to process natural language instructions and 3D scene data. This approach facilitates enhanced reasoning and response generation in complex 3D environments. Using the ScanNet 3D scene dataset, along with text annotations from ScanQA and ScanRefer, 3D-MoRe generates 62,000 question-answer (QA) pairs and 73,000 object descriptions across 1,513 scenes. We also employ various data augmentation techniques and implement semantic filtering to ensure high-quality data. Experiments on ScanQA demonstrate that 3D-MoRe significantly outperforms state-of-the-art baselines, with the CIDEr score improving by 2.15\%. Similarly, on ScanRefer, our approach achieves a notable increase in [email protected] by 1.84\%, highlighting its effectiveness in both tasks. Our code and generated datasets will be publicly released to benefit the community, and both can be accessed on the https://3D-MoRe.github.io.

3D问答多模态数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。