arXiv:2503.15208cs.CV2025-03ICCV被引 44

用度量深度解耦时空扩散,实现无需优化的4D驾驶场景生成。

DiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene Generation

  • 拆分时空生成为两个扩散过程,用度量深度统一建模。
  • 在真实数据集上实现领先的时间预测与新视角合成效果。
  • 适合自动驾驶场景模拟、规划系统训练等应用。

现有生成模型难以同时实现动态4D驾驶场景的时间外推与空间新视角合成(NVS),且无需每场景优化。核心挑战在于如何构建高效通用的几何表示以连接时空生成。为此,我们提出首个解耦的时空扩散框架DiST-4D,以度量深度为核心几何表示。该框架将问题分解为两个扩散过程:DiST-T直接从历史观测中预测未来度量深度与多视角RGB序列;DiST-S仅基于已有视角训练即可实现空间NVS,通过前向-反向渲染约束实现循环一致性,缩小观测与未见视角间的泛化差距。度量深度对精确预测与高保真空间合成至关重要,因其提供视图一致且可泛化的几何表征。实验表明,DiST-4D在时间预测与NVS任务上均达当前最优表现,并在规划相关评估中保持竞争力。

原文摘要 · Abstract (English)

Current generative models struggle to synthesize dynamic 4D driving scenes that simultaneously support temporal extrapolation and spatial novel view synthesis (NVS) without per-scene optimization. A key challenge lies in finding an efficient and generalizable geometric representation that seamlessly connects temporal and spatial synthesis. To address this, we propose DiST-4D, the first disentangled spatiotemporal diffusion framework for 4D driving scene generation, which leverages metric depth as the core geometric representation. DiST-4D decomposes the problem into two diffusion processes: DiST-T, which predicts future metric depth and multi-view RGB sequences directly from past observations, and DiST-S, which enables spatial NVS by training only on existing viewpoints while enforcing cycle consistency. This cycle consistency mechanism introduces a forward-backward rendering constraint, reducing the generalization gap between observed and unseen viewpoints. Metric depth is essential for both accurate reliable forecasting and accurate spatial NVS, as it provides a view-consistent geometric representation that generalizes well to unseen perspectives. Experiments demonstrate that DiST-4D achieves state-of-the-art performance in both temporal prediction and NVS tasks, while also delivering competitive performance in planning-related evaluations.

4D生成扩散模型自动驾驶新视角合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。