arXiv:2503.13185cs.CVcs.AI2025-03被引 10

用3D坐标轴提示让GPT-4o理解真实场景中的物体位置。

3DAxisPrompt: Promoting the 3D Grounding and Reasoning in GPT-4o

  • 通过SAM生成的掩码与3D坐标轴提供几何先验,增强模型3D感知能力。
  • 在ScanRefer、ScanNet等四个数据集上验证有效,显著提升3D定位性能。
  • 揭示了提示工程对3D视觉推理的潜力与局限,适合关注3D理解的研究者。

多模态大语言模型(MLLMs)在多种任务中表现出色,尤其在精心设计的视觉提示下。然而,现有研究主要集中在逻辑推理和2D视觉理解,而MLLM在3D视觉中的表现仍待探索。本文提出一种新型视觉提示方法3DAxisPrompt,旨在激发MLLM在真实场景中的3D理解能力。该方法利用分割任意模型(SAM)生成的掩码与3D坐标轴,为MLLM提供明确的几何先验,从而将其强大的2D定位与推理能力拓展至真实世界的3D场景。我们首次系统性地研究了多种视觉提示格式,并基于GPT-4o得出关于3D理解潜力与局限性的结论。此外,构建了包含ScanRefer、ScanNet、FMB和nuScene四个数据集的评估环境,开展了广泛的定量与定性实验,结果表明所提方法在多个3D任务中有效。总体而言,3DAxisPrompt使MLLM能有效感知物体在真实场景中的3D位置,但单一提示策略无法在所有3D任务中持续取得最优效果。本研究证实了通过提示工程实现MLLM在3D视觉定位/推理中的可行性。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) exhibit impressive capabilities across a variety of tasks, especially when equipped with carefully designed visual prompts. However, existing studies primarily focus on logical reasoning and visual understanding, while the capability of MLLMs to operate effectively in 3D vision remains an ongoing area of exploration. In this paper, we introduce a novel visual prompting method, called 3DAxisPrompt, to elicit the 3D understanding capabilities of MLLMs in real-world scenes. More specifically, our method leverages the 3D coordinate axis and masks generated from the Segment Anything Model (SAM) to provide explicit geometric priors to MLLMs and then extend their impressive 2D grounding and reasoning ability to real-world 3D scenarios. Besides, we first provide a thorough investigation of the potential visual prompting formats and conclude our findings to reveal the potential and limits of 3D understanding capabilities in GPT-4o, as a representative of MLLMs. Finally, we build evaluation environments with four datasets, i.e., ScanRefer, ScanNet, FMB, and nuScene datasets, covering various 3D tasks. Based on this, we conduct extensive quantitative and qualitative experiments, which demonstrate the effectiveness of the proposed method. Overall, our study reveals that MLLMs, with the help of 3DAxisPrompt, can effectively perceive an object's 3D position in real-world scenarios. Nevertheless, a single prompt engineering approach does not consistently achieve the best outcomes for all 3D tasks. This study highlights the feasibility of leveraging MLLMs for 3D vision grounding/reasoning with prompt engineering techniques.

3D理解提示工程多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。