arXiv:2510.07944cs.CV2025-10被引 2

跨视角视频生成模型,可同时输出高质量视频与深度信息。

CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving

  • 用时空重建VAE增强3D结构与动态编码能力
  • 在FID和FVD指标上显著优于现有方法
  • 适合自动驾驶场景模拟与几何理解任务

生成模型已广泛用于环境模拟与未来状态预测。随着自动驾驶发展,不仅需要在多种控制下生成高保真视频,还需产出多样且有意义的信息,如深度估计。为此,我们提出CVD-STORM,一种基于时空重建变分自编码器(VAE)的跨视角视频扩散模型,可在不同控制输入下生成具有4D重建能力的长时序多视角视频。首先,通过辅助4D重建任务微调VAE,提升其对三维结构与时间动态的编码能力;随后将该VAE融入视频扩散过程,显著改善生成质量。实验表明,模型在FID和FVD指标上均有显著提升。此外,联合训练的高斯点云解码器能有效重建动态场景,提供有助于全面场景理解的几何信息。

原文摘要 · Abstract (English)

Generative models have been widely applied to world modeling for environment simulation and future state prediction. With advancements in autonomous driving, there is a growing demand not only for high-fidelity video generation under various controls, but also for producing diverse and meaningful information such as depth estimation. To address this, we propose CVD-STORM, a cross-view video diffusion model utilizing a spatial-temporal reconstruction Variational Autoencoder (VAE) that generates long-term, multi-view videos with 4D reconstruction capabilities under various control inputs. Our approach first fine-tunes the VAE with an auxiliary 4D reconstruction task, enhancing its ability to encode 3D structures and temporal dynamics. Subsequently, we integrate this VAE into the video diffusion process to significantly improve generation quality. Experimental results demonstrate that our model achieves substantial improvements in both FID and FVD metrics. Additionally, the jointly-trained Gaussian Splatting Decoder effectively reconstructs dynamic scenes, providing valuable geometric information for comprehensive scene understanding. Our project page is https://sensetime-fvg.github.io/CVD-STORM.

视频生成自动驾驶扩散模型4D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。