综述稀疏视角三维重建新方法与挑战,助力机器人、AR/VR等场景
Sparse-View 3D Reconstruction: Recent Advances and Open Challenges
- 融合神经隐式、点云和生成模型,提升稀疏视角重建质量
- 对比实验揭示精度、效率与泛化间的关键权衡
- 适合关注3D生成、视觉定位与实时重建的研究者
稀疏视角三维重建对机器人、增强/虚拟现实(AR/VR)及自动驾驶等场景至关重要,因图像获取受限导致重叠度低,传统结构光从运动(SfM)与多视图立体(MVS)方法失效。本综述系统分析神经隐式模型(如NeRF及其正则化变体)、基于显式点云的方法(如3D Gaussian Splatting),以及利用扩散模型和视觉基础模型先验的混合框架。重点探讨几何正则化、显式形状建模与生成推理在缓解浮点物与位姿歧义方面的机制。标准基准测试表明,重建精度、效率与泛化能力存在显著权衡。不同于以往综述,本文统一梳理几何驱动、神经隐式与生成式(扩散基)方法。指出领域泛化与无位姿重建仍是核心挑战,并展望发展3D原生生成先验、实现实时无约束稀疏视角重建的未来方向。
原文摘要 · Abstract (English)
Sparse-view 3D reconstruction is essential for applications in which dense image acquisition is impractical, such as robotics, augmented/virtual reality (AR/VR), and autonomous systems. In these settings, minimal image overlap prevents reliable correspondence matching, causing traditional methods, such as structure-from-motion (SfM) and multiview stereo (MVS), to fail. This survey reviews the latest advances in neural implicit models (e.g., NeRF and its regularized versions), explicit point-cloud-based approaches (e.g., 3D Gaussian Splatting), and hybrid frameworks that leverage priors from diffusion and vision foundation models (VFMs).We analyze how geometric regularization, explicit shape modeling, and generative inference are used to mitigate artifacts such as floaters and pose ambiguities in sparse-view settings. Comparative results on standard benchmarks reveal key trade-offs between the reconstruction accuracy, efficiency, and generalization. Unlike previous reviews, our survey provides a unified perspective on geometry-based, neural implicit, and generative (diffusion-based) methods. We highlight the persistent challenges in domain generalization and pose-free reconstruction and outline future directions for developing 3D-native generative priors and achieving real-time, unconstrained sparse-view reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。