arXiv:2608.01588cs.CV2026-08

用双源深度先验,让稀疏摄像头也能高质量生成动态4D场景。

D^2-4DGS: Dual-Depth Guided Sparse-Camera 4D Gaussian Splatting

论文配图:D^2-4DGS: Dual-Depth Guided Sparse-Camera 4D Gaussian Splatting
图 1 · 摘自论文原文
  • 融合单目与多视角深度,自动筛选可靠几何锚点。
  • 在稀疏视图下仍保持结构完整,平均提升1.33 dB PSNR。
  • 适合低资源部署的动态三维重建,尤其适用于少相机场景。

动态4D高斯溅射已成为一种高效动态新视角合成方法,通过显式建模场景实现实时渲染。然而现有方法通常需要密集多视角视频以获得充分几何约束,导致采集成本高,限制了稀疏相机的应用。减少输入视角可降低获取成本,但削弱了几何监督,常导致结构缺失和漂浮高斯点。深度先验提供几何线索,但单一来源无法同时具备稠密覆盖与可靠几何。单目深度虽覆盖稠密但存在尺度模糊和局部偏差,而多视角几何深度则提供不完整的坐标一致锚点。为利用二者互补性,我们提出D²-4DGS,一种基于双源深度先验的稀疏相机动态4D高斯溅射框架。我们将单目估计与有效多视角几何深度对齐,并验证其一致性以识别可靠几何锚点。这些经验证的锚点支持一致性感知裁剪与深度监督,而验证后的几何深度与对齐的单目估计共同为重建不足区域提供候选几何以进行稠密化。最后,通过RGB-D联合优化,在稀疏视角监督下提升外观保真度与几何一致性。在全部九个数据集-视图设置中,D²-4DGS均达到最高PSNR,相比最优对比方法平均提升1.33 dB。

原文摘要 · Abstract (English)

Dynamic 4D Gaussian Splatting has emerged as an efficient representation for dynamic novel view synthesis through explicit scene modeling and real-time rendering. However, existing methods typically require dense multi-view videos for sufficient geometric constraints, making capture expensive and limiting sparse-camera deployment. Reducing input views lowers acquisition cost but weakens geometry supervision, often causing missing structures and floating Gaussians. Depth priors provide geometric cues, yet no single source offers both dense coverage and reliable geometry. Monocular depth provides dense structure but is scale-ambiguous and locally biased, whereas multi-view geometric depth provides incomplete anchors consistent with the reconstruction coordinate system. To exploit their complementarity, we propose D$^2$-4DGS, a sparse-camera dynamic 4D Gaussian Splatting framework guided by dual-source depth priors. We align monocular estimates with valid multi-view geometric depths and verify their consistency to identify reliable geometric anchors. These verified anchors support consistency-aware pruning and depth supervision, while verified geometric depths and aligned mono-only estimates provide candidate geometry for densification in under-reconstructed regions. Finally, RGB-D joint optimization improves appearance fidelity and geometric consistency under sparse-view supervision. Across all nine dataset--view settings, D$^2$-4DGS achieves the highest PSNR, improving by 1.33 dB on average over the best competing method in each setting.

4D高斯动态重建稀疏视角深度先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。