arXiv:2501.18594cs.CV2025-01综述被引 19

综述3D点云基础模型研究进展,探索多模态融合与语言增强的潜力。

Foundational Models for 3D Point Clouds: A Survey and Outlook

  • 系统梳理3D点云基础模型构建策略
  • 总结多模态融合在3D感知任务中的应用进展
  • 适合关注3D视觉与大模型交叉研究的读者

3D点云表示在保持物理世界几何保真度方面至关重要,有助于构建更精确的复杂3D环境。尽管人类通过多感官系统自然理解物体间关系与变化,人工智能系统尚未完全复制此能力。为弥合这一差距,需引入多模态信息。能无缝整合并跨模态推理的模型称为基础模型(FMs)。2D模态(如图像、文本)的基础模型发展迅速,得益于大规模数据集的支持。但3D领域因标注数据稀缺和高计算开销而进展缓慢。近年研究开始探索将2D知识迁移至3D任务,以克服挑战。同时,语言凭借其抽象推理与环境描述能力,为提升3D理解提供了新路径,特别是结合大型预训练语言模型(LLMs)。尽管近年来3D视觉任务中基础模型快速发展并被广泛应用,仍缺乏全面深入的文献综述。本文旨在填补该空白,系统回顾当前利用基础模型进行3D视觉理解的前沿方法。首先分析各类3D基础模型的构建策略;其次对不同基础模型在感知等任务中的应用进行分类总结;最后展望未来研究方向。附录提供相关论文清单:https://github.com/vgthengane/Awesome-FMs-in-3D。

原文摘要 · Abstract (English)

The 3D point cloud representation plays a crucial role in preserving the geometric fidelity of the physical world, enabling more accurate complex 3D environments. While humans naturally comprehend the intricate relationships between objects and variations through a multisensory system, artificial intelligence (AI) systems have yet to fully replicate this capacity. To bridge this gap, it becomes essential to incorporate multiple modalities. Models that can seamlessly integrate and reason across these modalities are known as foundation models (FMs). The development of FMs for 2D modalities, such as images and text, has seen significant progress, driven by the abundant availability of large-scale datasets. However, the 3D domain has lagged due to the scarcity of labelled data and high computational overheads. In response, recent research has begun to explore the potential of applying FMs to 3D tasks, overcoming these challenges by leveraging existing 2D knowledge. Additionally, language, with its capacity for abstract reasoning and description of the environment, offers a promising avenue for enhancing 3D understanding through large pre-trained language models (LLMs). Despite the rapid development and adoption of FMs for 3D vision tasks in recent years, there remains a gap in comprehensive and in-depth literature reviews. This article aims to address this gap by presenting a comprehensive overview of the state-of-the-art methods that utilize FMs for 3D visual understanding. We start by reviewing various strategies employed in the building of various 3D FMs. Then we categorize and summarize use of different FMs for tasks such as perception tasks. Finally, the article offers insights into future directions for research and development in this field. To help reader, we have curated list of relevant papers on the topic: https://github.com/vgthengane/Awesome-FMs-in-3D.

3D视觉基础模型多模态点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。