arXiv:2412.00493cs.CVcs.CL2024-12CVPR被引 157

将3D场景视为动态视频,提升大模型对空间位置的理解能力。

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

  • 把3D场景当作动态视频处理,加入位置编码增强空间感知。
  • 在多个基准上达到当前最优性能,如ScanRefer、ScanQA等。
  • 适合需要精准3D空间理解的机器人、自动驾驶研究者。

多模态大语言模型(MLLMs)在多种多模态任务中取得显著进展,但在需要理解3D环境空间结构的任务中仍面临挑战。现有方法虽尝试引入点云特征,但模型表示与3D场景的内在复杂性之间仍有较大差距,主要源于训练数据以2D为主,限制了其对3D空间的建模能力。为此,本文提出一种新型通用模型Video-3D LLM,用于3D场景理解。通过将3D场景视为动态视频,并在表征中融入3D位置编码,使视频表示更贴近真实空间上下文。此外,采用最大覆盖采样技术,在计算开销与性能间实现优化平衡。大量实验表明,该模型在多个3D场景理解基准测试中表现领先,包括ScanRefer、Multi3DRefer、Scan2Cap、ScanQA和SQA3D。

原文摘要 · Abstract (English)

The rapid advancement of Multimodal Large Language Models (MLLMs) has significantly impacted various multimodal tasks. However, these models face challenges in tasks that require spatial understanding within 3D environments. Efforts to enhance MLLMs, such as incorporating point cloud features, have been made, yet a considerable gap remains between the models' learned representations and the inherent complexity of 3D scenes. This discrepancy largely stems from the training of MLLMs on predominantly 2D data, which restricts their effectiveness in comprehending 3D spaces. To address this issue, in this paper, we propose a novel generalist model, i.e., Video-3D LLM, for 3D scene understanding. By treating 3D scenes as dynamic videos and incorporating 3D position encoding into these representations, our Video-3D LLM aligns video representations with real-world spatial contexts more accurately. In addition, we have implemented a maximum coverage sampling technique to optimize the trade-off between computational cost and performance. Extensive experiments demonstrate that our model achieves state-of-the-art performance on several 3D scene understanding benchmarks, including ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, and SQA3D.

3D理解视频建模位置编码多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。