生成20秒以上真实且几何一致的驾驶场景视频
AutoScape: Geometry-Consistent Long-Horizon Scene Generation
- 用扩散模型分步生成稀疏关键帧,保持几何一致性
- 生成20秒以上视频,长时序FID/FVD指标提升超40%
- 适合自动驾驶仿真与场景生成研究者
本文提出AutoScape,一种长时序驾驶场景生成框架。核心是新型的RGB-D扩散模型,通过迭代生成稀疏且几何一致的关键帧,作为场景外观与结构的可靠锚点。为保证长距离几何一致性,该模型在共享潜在空间中联合处理图像与深度图,显式依赖先前生成关键帧的渲染点云进行条件建模,并利用光流一致性引导采样过程。获得高质量的RGB-D关键帧后,视频扩散模型在它们之间插值生成稠密连贯的视频帧。AutoScape可生成超过20秒的真实驾驶视频,在长时序生成任务中,相较现有最优方法,FID与FVD指标分别提升48.6%和43.0%。
原文摘要 · Abstract (English)
This paper proposes AutoScape, a long-horizon driving scene generation framework. At its core is a novel RGB-D diffusion model that iteratively generates sparse, geometrically consistent keyframes, serving as reliable anchors for the scene's appearance and geometry. To maintain long-range geometric consistency, the model 1) jointly handles image and depth in a shared latent space, 2) explicitly conditions on the existing scene geometry (i.e., rendered point clouds) from previously generated keyframes, and 3) steers the sampling process with a warp-consistent guidance. Given high-quality RGB-D keyframes, a video diffusion model then interpolates between them to produce dense and coherent video frames. AutoScape generates realistic and geometrically consistent driving videos of over 20 seconds, improving the long-horizon FID and FVD scores over the prior state-of-the-art by 48.6\% and 43.0\%, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。