用扩散模型生成虚拟视角,解决大场景重建难题
Diffusion-Guided Gaussian Splatting for Large-Scale Unconstrained 3D Reconstruction and Novel View Synthesis
- 用多视角扩散模型生成伪观测数据,让稀疏输入也能重建
- 在四个基准上超越现有方法,尤其在复杂光照和遮挡下表现突出
- 适合做城市级3D重建的科研人员和工业应用开发者
3D高斯溅射(3DGS)和神经辐射场(NeRF)在实时3D重建与新视角合成方面取得显著进展。然而,在大规模、无约束环境中,由于输入覆盖稀疏不均、瞬时遮挡、外观变化大及相机参数不一致,性能严重下降。本文提出GS-Diff,一种由多视图扩散模型引导的3DGS框架,通过条件生成伪观测数据,将欠约束问题转化为良好定义的问题,实现稀疏数据下的鲁棒优化。该方法融合外观嵌入、单目深度先验、动态物体建模、各向异性正则化及先进光栅化技术,有效应对真实场景中的几何与光度挑战。在四个基准上的实验表明,GS-Diff持续大幅优于当前最优基线。
原文摘要 · Abstract (English)
Recent advancements in 3D Gaussian Splatting (3DGS) and Neural Radiance Fields (NeRF) have achieved impressive results in real-time 3D reconstruction and novel view synthesis. However, these methods struggle in large-scale, unconstrained environments where sparse and uneven input coverage, transient occlusions, appearance variability, and inconsistent camera settings lead to degraded quality. We propose GS-Diff, a novel 3DGS framework guided by a multi-view diffusion model to address these limitations. By generating pseudo-observations conditioned on multi-view inputs, our method transforms under-constrained 3D reconstruction problems into well-posed ones, enabling robust optimization even with sparse data. GS-Diff further integrates several enhancements, including appearance embedding, monocular depth priors, dynamic object modeling, anisotropy regularization, and advanced rasterization techniques, to tackle geometric and photometric challenges in real-world settings. Experiments on four benchmarks demonstrate that GS-Diff consistently outperforms state-of-the-art baselines by significant margins.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。