用扩散模型+高斯点云重建城市场景,视角变化时更稳定。
MuDG: Taming Multi-modal Diffusion with Gaussian Splatting for Urban Scene Reconstruction
- 结合多模态扩散与高斯点云,生成高质量新视角图像
- 在开放Waymo数据集上重建和合成效果优于现有方法
- 适合自动驾驶场景的3D重建与视觉一致性要求
最近辐射场的突破显著推进了自动驾驶中的3D场景重建与新视角合成(NVS)。然而,基于重建的方法在训练轨迹外大幅视角偏移时性能严重下降,而基于生成的方法则面临时间一致性差与场景控制精度不足的问题。为此,我们提出MuDG,一种将多模态扩散模型与高斯点云(GS)结合的城市场景重建框架。MuDG利用融合的LiDAR点云、RGB图像及几何先验来引导多模态视频扩散模型,生成逼真的新视角RGB、深度与语义输出。该合成流程支持前向NVS,无需耗时的逐场景优化,并提供全面监督信号以增强3DGS表示在极端视角下的渲染鲁棒性。在Open Waymo数据集上的实验表明,MuDG在重建与合成质量上均优于现有方法。
原文摘要 · Abstract (English)
Recent breakthroughs in radiance fields have significantly advanced 3D scene reconstruction and novel view synthesis (NVS) in autonomous driving. Nevertheless, critical limitations persist: reconstruction-based methods exhibit substantial performance deterioration under significant viewpoint deviations from training trajectories, while generation-based techniques struggle with temporal coherence and precise scene controllability. To overcome these challenges, we present MuDG, an innovative framework that integrates Multi-modal Diffusion model with Gaussian Splatting (GS) for Urban Scene Reconstruction. MuDG leverages aggregated LiDAR point clouds with RGB and geometric priors to condition a multi-modal video diffusion model, synthesizing photorealistic RGB, depth, and semantic outputs for novel viewpoints. This synthesis pipeline enables feed-forward NVS without computationally intensive per-scene optimization, providing comprehensive supervision signals to refine 3DGS representations for rendering robustness enhancement under extreme viewpoint changes. Experiments on the Open Waymo Dataset demonstrate that MuDG outperforms existing methods in both reconstruction and synthesis quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。