综述前馈式3D重建与视角合成最新进展,助力虚拟现实等应用。
Advances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey
- 按点云、3DGS、NeRF等表示架构分类,系统梳理前馈方法
- 涵盖无姿态重建、动态3D建模等关键任务,支持数字人等应用
- 总结常用数据集与评估标准,指明未来研究方向
3D重建与视角合成是计算机视觉、图形学及增强现实(AR)、虚拟现实(VR)、数字孪生等沉浸式技术的基础问题。传统方法依赖计算量大的迭代优化流程,限制了其在真实场景中的应用。近年来,深度学习驱动的前馈方法彻底改变了这一领域,实现了快速且泛化的3D重建与视角合成。本综述全面回顾了前馈式3D重建与视图合成技术,按底层表示架构(如点云、3D高斯泼溅3DGS、神经辐射场NeRF)进行分类。我们分析了无姿态重建、动态3D重建、3D感知图像与视频生成等关键任务,突出其在数字人、SLAM、机器人等领域的应用。此外,还梳理了常用数据集及其详细统计信息,以及各类下游任务的评估协议。最后讨论了开放性挑战与未来研究方向,强调前馈方法在推动3D视觉技术进步方面的潜力。
原文摘要 · Abstract (English)
3D reconstruction and view synthesis are foundational problems in computer vision, graphics, and immersive technologies such as augmented reality (AR), virtual reality (VR), and digital twins. Traditional methods rely on computationally intensive iterative optimization in a complex chain, limiting their applicability in real-world scenarios. Recent advances in feed-forward approaches, driven by deep learning, have revolutionized this field by enabling fast and generalizable 3D reconstruction and view synthesis. This survey offers a comprehensive review of feed-forward techniques for 3D reconstruction and view synthesis, with a taxonomy according to the underlying representation architectures including point cloud, 3D Gaussian Splatting (3DGS), Neural Radiance Fields (NeRF), etc. We examine key tasks such as pose-free reconstruction, dynamic 3D reconstruction, and 3D-aware image and video synthesis, highlighting their applications in digital humans, SLAM, robotics, and beyond. In addition, we review commonly used datasets with detailed statistics, along with evaluation protocols for various downstream tasks. We conclude by discussing open research challenges and promising directions for future work, emphasizing the potential of feed-forward approaches to advance the state of the art in 3D vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。