arXiv:2502.00342cs.CV2025-02中稿 · Information Fusion综述被引 15

综述3D场景问答任务,梳理数据集、方法与评估标准。

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering

  • 从数据集、方法、评测三方面系统梳理3D SQA研究
  • 发现多类方法共享架构模式,推动统一分析
  • 适合想了解3D理解前沿的AI研究人员

3D场景问答(3D SQA)是融合3D视觉感知与自然语言处理的跨学科任务,使智能体能够理解并交互复杂3D环境。大模型发展催生了多样化的数据集,并推动指令微调与零样本方法在3D SQA中的应用。然而,快速进展带来统一分析与比较的挑战。本文首次对3D SQA进行系统性综述,从数据集、方法与评估指标三方面组织现有工作。除基础分类外,识别出方法间的共性架构模式。进一步总结核心局限,探讨指令微调、多模态对齐与零样本等趋势对未来发展的启示。最后提出多个有前景的研究方向:数据集构建、任务泛化、交互建模与统一评估协议。本工作旨在为未来研究奠定基础,推动更通用、智能的3D SQA系统发展。

原文摘要 · Abstract (English)

3D Scene Question Answering (3D SQA) represents an interdisciplinary task that integrates 3D visual perception and natural language processing, empowering intelligent agents to comprehend and interact with complex 3D environments. Recent advances in large multimodal modelling have driven the creation of diverse datasets and spurred the development of instruction-tuning and zero-shot methods for 3D SQA. However, this rapid progress introduces challenges, particularly in achieving unified analysis and comparison across datasets and baselines. In this survey, we provide the first comprehensive and systematic review of 3D SQA. We organize existing work from three perspectives: datasets, methodologies, and evaluation metrics. Beyond basic categorization, we identify shared architectural patterns across methods. Our survey further synthesizes core limitations and discusses how current trends, such as instruction tuning, multimodal alignment, and zero-shot, can shape future developments. Finally, we propose a range of promising research directions covering dataset construction, task generalization, interaction modeling, and unified evaluation protocols. This work aims to serve as a foundation for future research and foster progress toward more generalizable and intelligent 3D SQA systems.

3D理解多模态问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。