arXiv:2511.16454cs.CV2025-11中稿 · AAAI被引 1

用多视角图像生成3D场景的立体表示,提升视觉语言模型理解能力

LLaVA$^3$: Representing 3D Scenes like a Cubist Painter to Boost 3D Scene Understanding of VLMs

  • 借鉴立体派绘画,将多视角信息融合成全景视觉表征
  • 仅用2D图像实现3D场景理解,无需额外微调
  • 在3D问答和语言定位任务中优于现有2D基线方法

由于缺乏3D训练数据,构建能够理解3D场景的多模态语言模型仍具挑战性,而视觉-语言模型(VLM)则依赖丰富的2D数据集。为此,我们提出LLaVA³(读作LLaVA-Cube),一种新方法,仅使用多视角2D图像即可提升VLM的3D场景理解能力,且无需任何微调。受立体派画家启发,该方法通过每个物体的全方位视觉表征来描述3D场景,这些表征来自场景的中间多视角3D重建结果。在3D VQA与3D语言定位任务上的大量实验表明,本方法显著优于以往基于2D图像的VLM方案。

原文摘要 · Abstract (English)

Developing a multi-modal language model capable of understanding 3D scenes remains challenging due to the limited availability of 3D training data, in contrast to the abundance of 2D datasets used for vision-language models (VLM). As an alternative, we introduce LLaVA$^3$ (pronounced LLaVA-Cube), a novel method that improves the 3D scene understanding capabilities of VLM using only multi-view 2D images and without any fine-tuning. Inspired by Cubist painters, who represented multiple viewpoints of a 3D object within a single picture, we propose to describe the 3D scene for the VLM through omnidirectional visual representations of each object. These representations are derived from an intermediate multi-view 3D reconstruction of the scene. Extensive experiments on 3D VQA and 3D language grounding show that our approach outperforms previous 2D-based VLM solutions.

3D理解视觉语言模型多视角重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。