无需SfM即可从极稀疏视角生成高质量3D图像
Enhancing Novel View Synthesis from extremely sparse views with SfM-free 3D Gaussian Splatting Framework
- 用密集立体匹配替代SfM,逐步估计相机位姿并重建全局稠密点云
- 仅需2个训练视图即实现PSNR提升2.75dB,显著优于现有方法
- 适合真实场景中视角稀疏的3D内容生成任务
3D Gaussian Splatting(3DGS)在新视角合成中表现出卓越的实时性能,但其效果高度依赖于具有精确相机位姿的密集多视角输入,而这类数据在真实场景中罕见。当输入视角极度稀疏时,3DGS所依赖的SfM初始化方法无法准确重建场景三维结构,导致渲染质量下降。本文提出一种无需SfM的3DGS方法,可联合估计相机位姿并重建3D场景。具体而言,我们引入一个密集立体模块,逐步估计相机位姿并重建全局稠密点云用于初始化。为解决极稀疏视角下的信息匮乏问题,提出一致视角插值模块,基于训练视图对插值相机位姿,并生成视角一致的内容作为训练监督信号。此外,引入多尺度拉普拉斯一致性正则化与自适应空间感知多尺度几何正则化,以提升几何结构与渲染质量。实验表明,本方法在极端稀疏视角条件下(仅使用2个训练视图)相比其他先进3DGS方法,PSNR提升达2.75dB。合成图像畸变极小,同时保留丰富高频细节,视觉质量显著优于现有技术。
原文摘要 · Abstract (English)
3D Gaussian Splatting (3DGS) has demonstrated remarkable real-time performance in novel view synthesis, yet its effectiveness relies heavily on dense multi-view inputs with precisely known camera poses, which are rarely available in real-world scenarios. When input views become extremely sparse, the Structure-from-Motion (SfM) method that 3DGS depends on for initialization fails to accurately reconstruct the 3D geometric structures of scenes, resulting in degraded rendering quality. In this paper, we propose a novel SfM-free 3DGS-based method that jointly estimates camera poses and reconstructs 3D scenes from extremely sparse-view inputs. Specifically, instead of SfM, we propose a dense stereo module to progressively estimates camera pose information and reconstructs a global dense point cloud for initialization. To address the inherent problem of information scarcity in extremely sparse-view settings, we propose a coherent view interpolation module that interpolates camera poses based on training view pairs and generates viewpoint-consistent content as additional supervision signals for training. Furthermore, we introduce multi-scale Laplacian consistent regularization and adaptive spatial-aware multi-scale geometry regularization to enhance the quality of geometrical structures and rendered content. Experiments show that our method significantly outperforms other state-of-the-art 3DGS-based approaches, achieving a remarkable 2.75dB improvement in PSNR under extremely sparse-view conditions (using only 2 training views). The images synthesized by our method exhibit minimal distortion while preserving rich high-frequency details, resulting in superior visual quality compared to existing techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。