arXiv:2603.16306cs.CV2026-03

解决自动驾驶场景重建中时空不一致问题,提升多视角画面一致性。

DriveFix: Spatio-Temporally Coherent Driving Scene Restoration

  • 采用交错扩散变压器架构,显式建模时序依赖与跨摄像头空间一致性。
  • 在Waymo、nuScenes等数据集上实现最佳重建与新视角合成效果。
  • 适合需要高精度4D场景重建的自动驾驶系统研发人员使用。

近期基于扩散先验的4D场景重建方法在自动驾驶新视角合成方面展现出潜力,但多数方法独立处理帧或按视图逐个处理,导致相机间空间错位和序列时间漂移。我们提出DriveFix,一种新型多视角修复框架,确保驾驶场景的时空一致性。该方法采用交错扩散变压器架构,配备专用模块以显式建模时序依赖和跨摄像头空间一致性。通过条件化历史上下文并引入几何感知训练损失,DriveFix强制恢复视图遵循统一3D几何结构,实现高质量纹理一致传播,显著减少伪影。在Waymo、nuScenes和PandaSet数据集上的广泛评估表明,DriveFix在重建与新视角合成任务中均达到当前最优性能,为真实世界部署中的鲁棒4D世界建模迈出关键一步。

原文摘要 · Abstract (English)

Recent advancements in 4D scene reconstruction, particularly those leveraging diffusion priors, have shown promise for novel view synthesis in autonomous driving. However, these methods often process frames independently or in a view-by-view manner, leading to a critical lack of spatio-temporal synergy. This results in spatial misalignment across cameras and temporal drift in sequences. We propose DriveFix, a novel multi-view restoration framework that ensures spatio-temporal coherence for driving scenes. Our approach employs an interleaved diffusion transformer architecture with specialized blocks to explicitly model both temporal dependencies and cross-camera spatial consistency. By conditioning the generation on historical context and integrating geometry-aware training losses, DriveFix enforces that the restored views adhere to a unified 3D geometry. This enables the consistent propagation of high-fidelity textures and significantly reduces artifacts. Extensive evaluations on the Waymo, nuScenes, and PandaSet datasets demonstrate that DriveFix achieves state-of-the-art performance in both reconstruction and novel view synthesis, marking a substantial step toward robust 4D world modeling for real-world deployment.

4D重建扩散模型自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。