arXiv:2604.03334cs.CV2026-04ICCV综述

梳理2D视觉模型适配3D分析的三大方法,揭示其权衡与前景。

Bridging the Dimensionality Gap: A Taxonomy and Survey of 2D Vision Model Adaptation for 3D Analysis

论文配图:Bridging the Dimensionality Gap: A Taxonomy and Survey of 2D Vision Model Adaptation for 3D Analysis
图 1 · 摘自论文原文
  • 按数据、架构、混合三类划分2D到3D的适配策略
  • 对比三类方法在计算开销、预训练依赖和几何先验保留上的差异
  • 适合关注3D视觉、多模态融合与基础模型的研究者

CNN与视觉变换器(ViT)在2D视觉中取得显著成功,推动了向复杂3D分析领域的扩展。然而,2D图像的规则密集网格与3D数据(如点云、网格)的不规则稀疏特性之间存在根本性差异。本文全面回顾并构建了一个统一的分类框架,将适配策略分为三类:(1) 数据中心方法,将3D数据投影至2D格式以利用现成2D模型;(2) 架构中心方法,设计原生3D网络;(3) 混合方法,协同结合两种范式,兼顾2D大规模数据的视觉先验与3D模型的显式几何推理。通过该框架,定性分析三类方法在计算复杂度、大规模预训练依赖及几何归纳偏置保留方面的核心权衡。讨论关键开放挑战,提出未来方向,包括3D基础模型、几何数据自监督学习(SSL)进展,以及多模态信号的深度整合。

原文摘要 · Abstract (English)

The remarkable success of Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) in 2D vision has spurred significant research in extending these architectures to the complex domain of 3D analysis. Yet, a core challenge arises from a fundamental dichotomy between the regular, dense grids of 2D images and the irregular, sparse nature of 3D data such as point clouds and meshes. This survey provides a comprehensive review and a unified taxonomy of adaptation strategies that bridge this gap, classifying them into three families: (1) Data-centric methods that project 3D data into 2D formats to leverage off-the-shelf 2D models, (2) Architecture-centric methods that design intrinsic 3D networks, and (3) Hybrid methods, which synergistically combine the two modeling paradigms to benefit from both rich visual priors of large 2D datasets and explicit geometric reasoning of 3D models. Through this framework, we qualitatively analyze the fundamental trade-offs between these families concerning computational complexity, reliance on large-scale pre-training, and the preservation of geometric inductive biases. We discuss key open challenges and outline promising future research directions, including the development of 3D foundation models, advancements in self-supervised learning (SSL) for geometric data, and the deeper integration of multi-modal signals.

3D分析模型适配视觉变换器点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。