arXiv:2606.28828cs.CV2026-06

用单目视频重建动态4D场景,兼顾几何一致性和逼真渲染。

Ground4D: Consistency-Aware 4D Reconstruction from Monocular Video

论文配图:Ground4D: Consistency-Aware 4D Reconstruction from Monocular Video
图 1 · 摘自论文原文
  • 先用3D基础模型初始化几何结构和相机位姿,保证多视角一致性。
  • 再通过可微渲染优化动态高斯点云,保持时空几何一致。
  • 支持任意时间点的视角合成,适合需要真实动态场景的应用。

从单目视频学习支持动态新视角合成且时间上保持精确几何的4D场景表示仍具挑战。动态高斯溅射通过光度优化实现强渲染效果,但未显式约束多视角几何一致性。相反,3D基础模型能恢复连贯的场景几何与相机运动,但其基于点的输出不适用于逼真渲染。我们提出Ground4D,一个分两阶段的几何引导框架:首先,通过无需训练的VGGT方法,从单目视频中重建多视角一致的3D几何与相机位姿,为动态高斯表示提供结构化可靠初始化;其次,通过动态高斯溅射进行几何一致性感知优化,利用可微渲染在观测与合成视角间维持多视图几何一致性。此外,Ground4D天然建模场景连续4D动态,自然支持任意时间戳的渲染。通过将基础级几何先验融入动态高斯优化,Ground4D在重建保真度与渲染性能上均取得提升,凸显几何引导约束在鲁棒4D建模中的关键作用。

原文摘要 · Abstract (English)

Learning a 4D scene representation from a single monocular video that supports dynamic novel-view synthesis while maintaining faithful geometry over time remains challenging. Dynamic Gaussian Splatting achieves strong rendering performance through photometric optimization, yet does not explicitly enforce multi-view geometric consistency. In contrast, 3D foundation models recover coherent scene geometry and camera motion, but their point-based outputs are not designed for photorealistic rendering. We propose Ground4D, a geometry-grounded framework built on two stages. First, we perform geometry initialization via 3D foundation models, leveraging VGGT in a training-free manner to reconstruct multi-view-consistent 3D geometry and camera poses from monocular video. The recovered geometry provides a structured and reliable initialization for dynamic Gaussian representations. Second, we conduct geometry-consistency-aware refinement via dynamic Gaussian Splatting, optimizing the representation through differentiable rendering while maintaining multi-view geometric consistency across both observed and synthesized viewpoints. Furthermore, Ground4D inherently models the continuous 4D dynamics of the scene, naturally supporting rendering at arbitrary timestamps. By integrating foundation-level geometric priors into dynamic Gaussian optimization, Ground4D achieves stronger reconstruction fidelity and rendering performance, underscoring the role of geometry-grounded constraints in robust 4D scene modeling.

4D重建动态高斯单目视频几何一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。