让大模型学会看懂3D场景,提升视觉理解能力
3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding
- 用预训练3D模型给2D大模型提供3D特征监督
- 多个任务上性能提升,尤其在跨视角理解上
- 适合做3D视觉理解的开发者和研究者
近期场景理解研究利用多模态大语言模型(MLLM)通过强大的2D预训练进行3D推理,但其预训练中缺乏显式3D数据,限制了3D表征能力。本文通过评估多视角对应关系,揭示3D表征质量与下游任务表现之间存在强正相关性。为此,我们提出3DRS框架,通过引入预训练3D基础模型的监督信号,对齐MLLM视觉特征与3D模型提取的丰富3D知识,有效增强其3D表征学习能力。在多个基准测试(包括视觉定位、图像描述生成和问答)及多种MLLM上进行的大量实验表明,该方法带来一致的性能提升。
原文摘要 · Abstract (English)
Recent advances in scene understanding have leveraged multimodal large language models (MLLMs) for 3D reasoning by capitalizing on their strong 2D pretraining. However, the lack of explicit 3D data during MLLM pretraining limits 3D representation capability. In this paper, we investigate the 3D-awareness of MLLMs by evaluating multi-view correspondence and reveal a strong positive correlation between the quality of 3D-aware representation and downstream task performance. Motivated by this, we propose 3DRS, a framework that enhances MLLM 3D representation learning by introducing supervision from pretrained 3D foundation models. Our approach aligns MLLM visual features with rich 3D knowledge distilled from 3D models, effectively improving scene understanding. Extensive experiments across multiple benchmarks and MLLMs -- including visual grounding, captioning, and question answering -- demonstrate consistent performance gains. Project page: https://visual-ai.github.io/3drs
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。