arXiv:2507.08448cs.CVcs.AI2025-07综述被引 12

用单次前向传播实现多视角3D重建,速度远超传统方法。

Review of Feed-forward 3D Reconstruction: From DUSt3R to VGGT

  • 基于Transformer的统一网络,一次推理完成相机位姿与三维结构估计。
  • 相比传统SfM/MVS,处理无纹理区域更鲁棒,计算效率提升数倍。
  • 适合快速部署于AR/VR、自动驾驶等实时场景,但动态物体仍存挑战。

3D重建旨在恢复场景的稠密三维结构,是增强现实、虚拟现实、自动驾驶和机器人等领域的重要技术基础。传统方法如运动恢复结构(SfM)和多视图立体(MVS)虽精度高,但依赖迭代优化,流程复杂、计算量大,且在无纹理区域表现差。近年来,深度学习推动了3D重建范式变革。以DUSt3R为代表的前馈模型,采用统一深度网络,在单次前向传播中直接从任意图像集联合推断相机位姿与稠密几何结构。本文系统综述该新兴领域:剖析其技术框架,包括基于Transformer的对应建模、位姿与几何联合回归机制,以及从双视图扩展到多视图的策略;对比传统流程与早期学习方法(如MVSNet)的差异;梳理相关数据集与评估指标;最后探讨其应用前景,并指出关键挑战,如模型精度、可扩展性及动态场景处理能力。

原文摘要 · Abstract (English)

3D reconstruction, which aims to recover the dense three-dimensional structure of a scene, is a cornerstone technology for numerous applications, including augmented/virtual reality, autonomous driving, and robotics. While traditional pipelines like Structure from Motion (SfM) and Multi-View Stereo (MVS) achieve high precision through iterative optimization, they are limited by complex workflows, high computational cost, and poor robustness in challenging scenarios like texture-less regions. Recently, deep learning has catalyzed a paradigm shift in 3D reconstruction. A new family of models, exemplified by DUSt3R, has pioneered a feed-forward approach. These models employ a unified deep network to jointly infer camera poses and dense geometry directly from an Unconstrained set of images in a single forward pass. This survey provides a systematic review of this emerging domain. We begin by dissecting the technical framework of these feed-forward models, including their Transformer-based correspondence modeling, joint pose and geometry regression mechanisms, and strategies for scaling from two-view to multi-view scenarios. To highlight the disruptive nature of this new paradigm, we contrast it with both traditional pipelines and earlier learning-based methods like MVSNet. Furthermore, we provide an overview of relevant datasets and evaluation metrics. Finally, we discuss the technology's broad application prospects and identify key future challenges and opportunities, such as model accuracy and scalability, and handling dynamic scenes.

3D重建前馈模型TransformerAR/VR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。