提出STAGE模型,实现长时间驾驶场景的高质量视频生成
STAGE: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation
- 采用分层时序特征传递机制,分离时序与去噪过程提升帧间一致性
- 多阶段训练策略加速收敛,600帧长视频生成远超现有方法极限
- 适合自动驾驶仿真、长序列生成任务的研究者和工程师
在自动驾驶世界建模中,生成长时间跨度、高保真度的驾驶视频面临重大挑战。现有方法因时空动态解耦不足和跨帧特征传播机制有限,常出现误差累积与特征错位问题。为此,我们提出STAGE(Streaming Temporal Attention Generative Engine),一种新颖的自回归框架,首次引入分层特征协调与多阶段优化,以实现可持续的视频合成。为提升长时视频生成质量,我们设计了分层时序特征传递(HTFT)机制,通过分离建模时序演化与去噪过程,并在帧间传递去噪特征,增强生成视频的时序一致性。同时,提出多阶段训练策略,将训练分为三个阶段,通过模型解耦与自回归推理过程模拟,有效加速模型收敛并减少误差累积。在Nuscenes数据集上的实验表明,STAGE在长时驾驶视频生成任务中显著优于现有方法。此外,我们探索了其生成无限长度驾驶视频的能力,成功生成600帧高质量视频,远超现有方法的最长生成长度。
原文摘要 · Abstract (English)
The generation of temporally consistent, high-fidelity driving videos over extended horizons presents a fundamental challenge in autonomous driving world modeling. Existing approaches often suffer from error accumulation and feature misalignment due to inadequate decoupling of spatio-temporal dynamics and limited cross-frame feature propagation mechanisms. To address these limitations, we present STAGE (Streaming Temporal Attention Generative Engine), a novel auto-regressive framework that pioneers hierarchical feature coordination and multi-phase optimization for sustainable video synthesis. To achieve high-quality long-horizon driving video generation, we introduce Hierarchical Temporal Feature Transfer (HTFT) and a novel multi-stage training strategy. HTFT enhances temporal consistency between video frames throughout the video generation process by modeling the temporal and denoising process separately and transferring denoising features between frames. The multi-stage training strategy is to divide the training into three stages, through model decoupling and auto-regressive inference process simulation, thereby accelerating model convergence and reducing error accumulation. Experiments on the Nuscenes dataset show that STAGE has significantly surpassed existing methods in the long-horizon driving video generation task. In addition, we also explored STAGE's ability to generate unlimited-length driving videos. We generated 600 frames of high-quality driving videos on the Nuscenes dataset, which far exceeds the maximum length achievable by existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。