解决稀疏视角下动态场景渲染几何不一致问题,实现高效高质实时渲染。
Geometry-Consistent 4D Gaussian Splatting for Sparse-Input Dynamic View Synthesis
- 引入时空一致性检查与全局-局部深度正则化,提升稀疏输入下的几何学习一致性。
- 在N3DV和Technicolor数据集上相较原4DGS提升1.58dB PSNR,优于RF-DeRF 2.62dB。
- 可部署于资源受限的物联网边缘设备,适合数字孪生等实际应用。
高斯点阵已被视为动态场景视图合成的新方法,在AIoT应用如数字孪生中展现巨大潜力。然而,现有动态高斯点阵方法在仅提供稀疏输入视角时性能显著下降,根源在于随视角减少导致的4D几何学习不一致。本文提出GC-4DGS框架,通过注入几何一致性,实现从稀疏视角出发的实时高质量动态场景渲染。尽管基于学习的多视角立体(MVS)与单目深度估计(MDE)可提供几何先验,但直接融合至4DGS因稀疏输入下4D几何优化病态性而效果不佳。为此,我们设计动态一致性检查策略以降低MVS在时空上的估计不确定性;并提出全局-局部深度正则化方法,从单目深度中提炼时空一致的几何信息,增强4D体内的几何与外观学习一致性。在主流N3DV与Technicolor数据集上的大量实验验证了该方法的有效性:相比原始4DGS与最新针对稀疏输入的动态辐射场RF-DeRF,PSNR分别提升1.58dB与2.62dB,且可在资源受限的物联网边缘设备上无缝部署。
原文摘要 · Abstract (English)
Gaussian Splatting has been considered as a novel way for view synthesis of dynamic scenes, which shows great potential in AIoT applications such as digital twins. However, recent dynamic Gaussian Splatting methods significantly degrade when only sparse input views are available, limiting their applicability in practice. The issue arises from the incoherent learning of 4D geometry as input views decrease. This paper presents GC-4DGS, a novel framework that infuses geometric consistency into 4D Gaussian Splatting (4DGS), offering real-time and high-quality dynamic scene rendering from sparse input views. While learning-based Multi-View Stereo (MVS) and monocular depth estimators (MDEs) provide geometry priors, directly integrating these with 4DGS yields suboptimal results due to the ill-posed nature of sparse-input 4D geometric optimization. To address these problems, we introduce a dynamic consistency checking strategy to reduce estimation uncertainties of MVS across spacetime. Furthermore, we propose a global-local depth regularization approach to distill spatiotemporal-consistent geometric information from monocular depths, thereby enhancing the coherent geometry and appearance learning within the 4D volume. Extensive experiments on the popular N3DV and Technicolor datasets validate the effectiveness of GC-4DGS in rendering quality without sacrificing efficiency. Notably, our method outperforms RF-DeRF, the latest dynamic radiance field tailored for sparse-input dynamic view synthesis, and the original 4DGS by 2.62dB and 1.58dB in PSNR, respectively, with seamless deployability on resource-constrained IoT edge devices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。