arXiv:2504.05786cs.CVcs.AI2025-04IJCAI综述被引 57

让大模型理解三维空间,突破传统视觉方法局限。

How to Enable LLM with 3D Capacity? A Survey of Spatial Reasoning in LLM

  • 分三类:从2D图像、点云数据或混合模态中提取三维信息
  • 提出统一分类框架,涵盖数据表示与跨模态训练策略
  • 适合关注机器人、自动驾驶等三维应用的研究者

三维空间理解在机器人、自动驾驶、虚拟现实和医学影像等实际应用中至关重要。近年来,大型语言模型(LLMs)在多个领域表现出色,被用于提升三维理解任务,展现出超越传统计算机视觉方法的潜力。本文全面综述了将LLM与三维空间理解相结合的方法。我们提出一个分类体系,将现有方法分为三类:基于图像的方法从二维视觉数据中推导三维理解;基于点云的方法直接处理三维表示;混合模态方法结合多种数据流。系统回顾了各类代表性方法,涵盖数据表示、架构改进及文本与三维模态间的训练策略。最后,讨论当前挑战,如数据集稀缺与计算负担,并指出空间感知、多模态融合和真实场景应用等有前景的研究方向。

原文摘要 · Abstract (English)

3D spatial understanding is essential in real-world applications such as robotics, autonomous vehicles, virtual reality, and medical imaging. Recently, Large Language Models (LLMs), having demonstrated remarkable success across various domains, have been leveraged to enhance 3D understanding tasks, showing potential to surpass traditional computer vision methods. In this survey, we present a comprehensive review of methods integrating LLMs with 3D spatial understanding. We propose a taxonomy that categorizes existing methods into three branches: image-based methods deriving 3D understanding from 2D visual data, point cloud-based methods working directly with 3D representations, and hybrid modality-based methods combining multiple data streams. We systematically review representative methods along these categories, covering data representations, architectural modifications, and training strategies that bridge textual and 3D modalities. Finally, we discuss current limitations, including dataset scarcity and computational challenges, while highlighting promising research directions in spatial perception, multi-modal fusion, and real-world applications.

三维理解大模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。