用视频生成模型实现可控制的4D驾驶场景模拟,提升真实感和视角自由度。
Stag-1: Towards Realistic 4D Driving Simulation with Video Generation Model
- 通过解耦时空关系,生成连贯的关键帧视频。
- 支持任意视角生成,多视角一致性与背景连贯性显著提升。
- 适合自动驾驶仿真、视觉算法验证等研究者使用。
4D驾驶仿真对开发逼真的自动驾驶模拟器至关重要。尽管现有方法在生成驾驶场景方面取得进展,但在视角变换和时空动态建模方面仍存在挑战。为此,我们提出空间-时间驱动仿真模型Stag-1,利用车载环视数据重建真实世界场景,并设计可控生成网络实现4D仿真。Stag-1基于多视角数据构建连续4D点云场景,解耦时空关系,生成一致的关键帧视频。同时,借助视频生成模型,从任意视角生成照片级真实感且可控的4D驾驶视频。为扩展视角生成范围,我们基于分解的相机位姿训练车辆运动视频,增强远距离场景建模能力。此外,通过重构车辆相机轨迹,融合连续视角的3D点云,实现时间维度上的全面场景理解。经多层次场景训练后,Stag-1可实现任意视角模拟,并在静态时空条件下深入理解场景演化。相比现有方法,本方案在多视角一致性、背景连贯性和准确性方面表现优异,推动了真实自动驾驶仿真技术的发展。代码已开源。
原文摘要 · Abstract (English)
4D driving simulation is essential for developing realistic autonomous driving simulators. Despite advancements in existing methods for generating driving scenes, significant challenges remain in view transformation and spatial-temporal dynamic modeling. To address these limitations, we propose a Spatial-Temporal simulAtion for drivinG (Stag-1) model to reconstruct real-world scenes and design a controllable generative network to achieve 4D simulation. Stag-1 constructs continuous 4D point cloud scenes using surround-view data from autonomous vehicles. It decouples spatial-temporal relationships and produces coherent keyframe videos. Additionally, Stag-1 leverages video generation models to obtain photo-realistic and controllable 4D driving simulation videos from any perspective. To expand the range of view generation, we train vehicle motion videos based on decomposed camera poses, enhancing modeling capabilities for distant scenes. Furthermore, we reconstruct vehicle camera trajectories to integrate 3D points across consecutive views, enabling comprehensive scene understanding along the temporal dimension. Following extensive multi-level scene training, Stag-1 can simulate from any desired viewpoint and achieve a deep understanding of scene evolution under static spatial-temporal conditions. Compared to existing methods, our approach shows promising performance in multi-view scene consistency, background coherence, and accuracy, and contributes to the ongoing advancements in realistic autonomous driving simulation. Code: https://github.com/wzzheng/Stag.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。