用扩散模型提升城市场景重建的视角泛化能力,几分钟内完成优化。
Diffusion-guided Generalizable Enhancer for Urban Scene Reconstruction

- 基于扩散模型学习跨场景生成先验,增强3D高斯表示。
- 在挑战性视角下(如变道)保持高质量,优于现有方法。
- 无需每场景重训练,适合自动驾驶仿真等大规模应用。
从真实世界观测中重建城市场景已成为自动驾驶开发与测试的重要工具。尽管当前神经渲染方法在记录轨迹上能实现高保真渲染,但在大幅视角变化时质量显著下降,限制了闭环仿真的应用。近期工作表明,利用扩散模型可提升这些困难视角的质量,并将改进结果回蒸至3D表示。然而,这些方法通常需要昂贵的每场景优化,且蒸馏后的表示仍脆弱,难以泛化到有限合成视图之外。为此,我们提出GenRe——一种新的扩散引导通用增强器,用于城市场景重建。GenRe以任意预训练3D高斯表示为输入,在几分钟内修复其缺陷。通过学习跨多样化场景的生成先验,GenRe高效生成鲁棒且高保真的表示,能可靠泛化至未见挑战性视角(如车道变更)。实验表明,GenRe在质量和效率上均优于现有方法,并提升多种下游任务表现,支持自动驾驶中稳健且可扩展的传感器仿真。
原文摘要 · Abstract (English)
Urban scene reconstruction from real-world observations has emerged as a powerful tool for self-driving development and testing. While current neural rendering approaches achieve high-fidelity rendering along the recorded trajectories, their quality degrades significantly under large viewpoint shifts, limiting the applicability for closed-loop simulation. Recent works have shown promising results in using diffusion models to enhance quality at these challenging viewpoints and distill improvements back into 3D representations. However, they often require costly per-scene optimization, and the distilled representations remain fragile and fail to generalize beyond limited synthesized views. To address these limitations, we propose GenRe, a novel diffusion-guided generalizable enhancer for urban scene reconstruction. GenRe takes as input any pretrained 3D Gaussian representation and fixes the deficiencies within a few minutes. By learning to distill generative priors across diverse scenes, GenRe produces robust and high-fidelity representation efficiently that generalizes reliably to challenging unseen viewpoints (e.g., lane change). Experiments show that GenRe outperforms existing methods in both quality and efficiency and benefits various downstream tasks, enabling robust and scalable sensor simulation for autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。