arXiv:2501.01428cs.CV2025-01被引 115

用视频+鸟瞰图提升视觉模型对3D场景的理解能力

GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models

  • 通过鸟瞰图与帧间物体标记建立全局-局部对应关系
  • 零样本下超越GPT-4o,16.5万条标注数据支持训练
  • 训练后模型推理时无需提示也能持续提升3D理解

近年来,2D视觉语言模型在图像-文本理解任务中取得显著进展,但在涉及具身智能的关键3D空间理解方面表现有限。现有方法依赖3D点云或多视角图像输入,而本文受人类感知启发,提出纯视觉解决方案。实证发现,当前模型在3D空间知识上的主要瓶颈在于场景与单帧之间的全局-局部对应缺失。为此,我们提出GPT4Scene,一种新的视觉提示范式,在训练和推理中构建鸟瞰图(BEV)与视频帧间的对应关系:从视频生成BEV图,并在帧与BEV中统一标记物体ID,将拼接后的BEV与视频帧共同输入模型。在零样本评估中,GPT4Scene性能超越闭源模型GPT-4o。此外,我们构建了包含16.5万条文本标注的处理后视频数据集,用于微调开源视觉语言模型,在所有3D理解任务上达到领先水平。令人惊讶的是,经该范式训练后,模型在推理阶段即使不使用显式提示或BEV图,性能仍持续提升,表明其内生具备了3D场景理解能力,为预训练模型无缝扩展至3D理解提供了新路径。

原文摘要 · Abstract (English)

In recent years, 2D Vision-Language Models (VLMs) have made significant strides in image-text understanding tasks. However, their performance in 3D spatial comprehension, which is critical for embodied intelligence, remains limited. Recent advances have leveraged 3D point clouds and multi-view images as inputs, yielding promising results. However, we propose exploring a purely vision-based solution inspired by human perception, which merely relies on visual cues for 3D spatial understanding. This paper empirically investigates the limitations of VLMs in 3D spatial knowledge, revealing that their primary shortcoming lies in the lack of global-local correspondence between the scene and individual frames. To address this, we introduce GPT4Scene, a novel visual prompting paradigm in VLM training and inference that helps build the global-local relationship, significantly improving the 3D spatial understanding of indoor scenes. Specifically, GPT4Scene constructs a Bird's Eye View (BEV) image from the video and marks consistent object IDs across both frames and the BEV image. The model then inputs the concatenated BEV image and video frames with markers. In zero-shot evaluations, GPT4Scene improves performance over closed-source VLMs like GPT-4o. Additionally, we prepare a processed video dataset consisting of 165K text annotation to fine-tune open-source VLMs, achieving state-of-the-art performance on all 3D understanding tasks. Surprisingly, after training with the GPT4Scene paradigm, VLMs consistently improve during inference, even without object marker prompting and BEV image as explicit correspondence. It demonstrates that the proposed paradigm helps VLMs develop an intrinsic ability to understand 3D scenes, which paves the way for a seamless approach to extending pre-trained VLMs for 3D scene understanding.

3D理解视觉语言模型鸟瞰图视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。