用多视角图像补足3D点云信息,提升大模型对场景的理解能力
Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models
- 融合多视角图像与相机位姿生成场景特征,增强3D理解
- 在多个下游任务中超越现有3D大模型表现
- 适合需要精细3D场景理解的视觉-语言应用
基础模型的发展使各类下游任务成为可能,尤其是大型语言模型(LLMs)在3D场景理解任务中展现出显著潜力。当前方法主要依赖3D点云,但室内场景重建常导致信息丢失,纹理缺失平面或重复图案易被遗漏,形成空洞;复杂结构物体则因图像与点云配准偏差引入细节失真。2D多视角图像与3D点云具有视觉一致性,能提供更丰富的场景细节,可自然弥补上述缺陷。为此,我们提出Argus,一种新型3D多模态框架,利用多视角图像增强大模型对3D场景的理解。Argus可视为3D大模态基础模型(3D-LMM),输入包括文本指令、2D多视角图像和3D点云,通过将多视角图像与相机位姿融合为视图-场景特征,并与3D特征交互,生成全面且细致的3D感知场景嵌入。该方法有效补偿点云重建中的信息损失,帮助LLMs更好理解3D世界。大量实验证明,本方法在多个下游任务中优于现有3D-LMMs。
原文摘要 · Abstract (English)
Advancements in foundation models have made it possible to conduct applications in various downstream tasks. Especially, the new era has witnessed a remarkable capability to extend Large Language Models (LLMs) for tackling tasks of 3D scene understanding. Current methods rely heavily on 3D point clouds, but the 3D point cloud reconstruction of an indoor scene often results in information loss. Some textureless planes or repetitive patterns are prone to omission and manifest as voids within the reconstructed 3D point clouds. Besides, objects with complex structures tend to introduce distortion of details caused by misalignments between the captured images and the dense reconstructed point clouds. 2D multi-view images present visual consistency with 3D point clouds and provide more detailed representations of scene components, which can naturally compensate for these deficiencies. Based on these insights, we propose Argus, a novel 3D multimodal framework that leverages multi-view images for enhanced 3D scene understanding with LLMs. In general, Argus can be treated as a 3D Large Multimodal Foundation Model (3D-LMM) since it takes various modalities as input(text instructions, 2D multi-view images, and 3D point clouds) and expands the capability of LLMs to tackle 3D tasks. Argus involves fusing and integrating multi-view images and camera poses into view-as-scene features, which interact with the 3D features to create comprehensive and detailed 3D-aware scene embeddings. Our approach compensates for the information loss while reconstructing 3D point clouds and helps LLMs better understand the 3D world. Extensive experiments demonstrate that our method outperforms existing 3D-LMMs in various downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。