用视频扩散模型提升动态城市场景的4D高斯点云重建效果
VDEGaussian: Video Diffusion Enhanced 4D Gaussian Splatting for Dynamic Urban Scenes Modeling
- 引入测试时自适应的视频扩散模型,提取时序一致的先验信息
- 通过帧位姿联合优化与不确定性蒸馏,提升快速运动物体建模精度
- 特别适合需要高保真动态重建的自动驾驶、虚拟现实应用
动态城市场景建模是快速发展且应用广泛的领域。尽管现有基于神经辐射场或高斯点云的方法已实现精细重建和高保真新视角合成,但仍面临依赖预标定物体轨迹、难以准确建模欠采样捕获下的快速运动物体等问题,尤其在处理时序不连续性时表现不足。为此,我们提出一种新型视频扩散增强的4D高斯点云框架。核心思路是从测试时自适应的视频扩散模型中蒸馏出鲁棒、时序一致的先验。为确保姿态精确对齐并有效融合去噪内容,提出两项创新:联合时间戳优化策略以精炼插值帧姿态,以及不确定性蒸馏方法以自适应提取目标内容并保留重建良好的区域。大量实验表明,该方法显著提升了动态建模能力,尤其在快速运动物体上表现优异,相比基线方法在新视角合成上获得约2 dB的PSNR提升。
原文摘要 · Abstract (English)
Dynamic urban scene modeling is a rapidly evolving area with broad applications. While current approaches leveraging neural radiance fields or Gaussian Splatting have achieved fine-grained reconstruction and high-fidelity novel view synthesis, they still face significant limitations. These often stem from a dependence on pre-calibrated object tracks or difficulties in accurately modeling fast-moving objects from undersampled capture, particularly due to challenges in handling temporal discontinuities. To overcome these issues, we propose a novel video diffusion-enhanced 4D Gaussian Splatting framework. Our key insight is to distill robust, temporally consistent priors from a test-time adapted video diffusion model. To ensure precise pose alignment and effective integration of this denoised content, we introduce two core innovations: a joint timestamp optimization strategy that refines interpolated frame poses, and an uncertainty distillation method that adaptively extracts target content while preserving well-reconstructed regions. Extensive experiments demonstrate that our method significantly enhances dynamic modeling, especially for fast-moving objects, achieving an approximate PSNR gain of 2 dB for novel view synthesis over baseline approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。