arXiv:2509.20251cs.CV2025-09被引 11

用立体强制策略生成动态4D驾驶场景,支持时空连续重建与新视角合成。

4D Driving Scene Generation With Stereo Forcing

  • 融合视频生成与几何一致性,构建统一的4D场景生成框架。
  • 在KITTI、Lyft HD等数据集上实现最优的外观与几何重建精度。
  • 适合自动驾驶仿真、虚拟驾驶训练等需要高保真4D场景的场景。

现有生成模型难以同时实现动态4D驾驶场景的时序外推与空间新视角合成(NVS),且无需每场景优化。本文提出PhiGenesis,一个统一的4D场景生成框架,扩展视频生成技术以保证几何与时间一致性。给定多视角图像序列和相机参数,该方法生成沿目标3D轨迹的时序连续4D高斯点云表示。第一阶段利用预训练视频VAE与新颖的范围视图适配器,实现从多视角图像的前向4D重建,支持单帧或视频输入,输出包含几何、语义与运动的完整4D场景。第二阶段引入几何引导的视频扩散模型,以渲染的历史4D场景为先验,根据轨迹生成未来视角。为缓解新视角中的几何偏差,提出立体强制(Stereo Forcing)条件策略,在去噪过程中融合几何不确定性,通过不确定性感知扰动动态调整生成影响。实验表明,该方法在外观与几何重建、时序生成及新视角合成任务中均达到当前最优性能,并在下游评估中表现优异。

原文摘要 · Abstract (English)

Current generative models struggle to synthesize dynamic 4D driving scenes that simultaneously support temporal extrapolation and spatial novel view synthesis (NVS) without per-scene optimization. Bridging generation and novel view synthesis remains a major challenge. We present PhiGenesis, a unified framework for 4D scene generation that extends video generation techniques with geometric and temporal consistency. Given multi-view image sequences and camera parameters, PhiGenesis produces temporally continuous 4D Gaussian splatting representations along target 3D trajectories. In its first stage, PhiGenesis leverages a pre-trained video VAE with a novel range-view adapter to enable feed-forward 4D reconstruction from multi-view images. This architecture supports single-frame or video inputs and outputs complete 4D scenes including geometry, semantics, and motion. In the second stage, PhiGenesis introduces a geometric-guided video diffusion model, using rendered historical 4D scenes as priors to generate future views conditioned on trajectories. To address geometric exposure bias in novel views, we propose Stereo Forcing, a novel conditioning strategy that integrates geometric uncertainty during denoising. This method enhances temporal coherence by dynamically adjusting generative influence based on uncertainty-aware perturbations. Our experimental results demonstrate that our method achieves state-of-the-art performance in both appearance and geometric reconstruction, temporal generation and novel view synthesis (NVS) tasks, while simultaneously delivering competitive performance in downstream evaluations. Homepage is at \href{https://jiangxb98.github.io/PhiGensis}{PhiGensis}.

4D生成驾驶场景扩散模型新视角合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。