arXiv:2503.23024cs.CV2025-03CVPR被引 7

让大模型理解3D场景中的视角变化,自动生成带位置信息的描述数据。

Empowering Large Language Models with 3D Situation Awareness

论文配图:Empowering Large Language Models with 3D Situation Awareness
图 1 · 摘自论文原文
  • 利用扫描轨迹和视觉语言模型自动生成带视角信息的数据。
  • 引入定位模块让模型能准确判断观察者的位置与朝向。
  • 显著提升大模型在3D场景中的空间描述能力,适合机器人导航等应用。

受大语言模型在2D图像领域成功启发,其在3D场景理解中的应用成为新趋势。3D与2D的关键区别在于:自我中心观察者的处境会变化,导致描述不同(如“左”或“右”)。然而,现有基于大语言模型的方法忽视了这种自我中心视角,仅使用全局视角的数据集。为此,我们提出一种新方法,通过利用数据采集过程中的扫描轨迹,并借助视觉-语言模型(VLMs)生成高质量的描述与问答对,自动构建具有情境感知能力的数据集。此外,我们引入一个情境定位模块,显式预测观察者视角的位置与方向,使大语言模型能够在3D场景中正确锚定情境描述。我们在多个基准上评估该方法,结果表明,该方法有效提升了大语言模型的3D情境感知能力,同时显著扩展了现有数据集并减少了人工工作量。

原文摘要 · Abstract (English)

Driven by the great success of Large Language Models (LLMs) in the 2D image domain, their applications in 3D scene understanding has emerged as a new trend. A key difference between 3D and 2D is that the situation of an egocentric observer in 3D scenes can change, resulting in different descriptions (e.g., ''left" or ''right"). However, current LLM-based methods overlook the egocentric perspective and simply use datasets from a global viewpoint. To address this issue, we propose a novel approach to automatically generate a situation-aware dataset by leveraging the scanning trajectory during data collection and utilizing Vision-Language Models (VLMs) to produce high-quality captions and question-answer pairs. Furthermore, we introduce a situation grounding module to explicitly predict the position and orientation of observer's viewpoint, thereby enabling LLMs to ground situation description in 3D scenes. We evaluate our approach on several benchmarks, demonstrating that our method effectively enhances the 3D situational awareness of LLMs while significantly expanding existing datasets and reducing manual effort.

3D理解大模型视角感知数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。