用运动感知注意力生成200帧以上的多视角驾驶视频。
DreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Scenes
- 基于扩散模型的自回归框架,融合透视引导和物体位置编码。
- 在16帧评估中质量优于基线,可生成超200帧长视频。
- 适配真实仿真器DriveArena,支持视觉驾驶代理测试。
扩散模型的进展提升了可控街景生成能力,并支持下游感知与规划任务。然而,在准确建模驾驶场景及生成长视频方面仍存在挑战。为此,我们提出DreamForge,一种专为3D可控长期生成设计的扩散自回归视频生成模型。为增强车道与前景生成,引入透视引导并集成对象级位置编码,以捕捉局部3D相关性并改进前景建模。同时提出运动感知时间注意力机制,以捕获视频中的运动线索与外观变化。通过利用运动帧与自回归生成范式,可在仅训练短序列的情况下,自回归生成超过200帧的长视频,在16帧视频评估中表现优于基线。最后,将方法集成至真实仿真器DriveArena,为基于视觉的驾驶智能体提供更可靠的开环与闭环评估。
原文摘要 · Abstract (English)
Recent advances in diffusion models have improved controllable streetscape generation and supported downstream perception and planning tasks. However, challenges remain in accurately modeling driving scenes and generating long videos. To alleviate these issues, we propose DreamForge, an advanced diffusion-based autoregressive video generation model tailored for 3D-controllable long-term generation. To enhance the lane and foreground generation, we introduce perspective guidance and integrate object-wise position encoding to incorporate local 3D correlation and improve foreground object modeling. We also propose motion-aware temporal attention to capture motion cues and appearance changes in videos. By leveraging motion frames and an autoregressive generation paradigm,we can autoregressively generate long videos (over 200 frames) using a model trained in short sequences, achieving superior quality compared to the baseline in 16-frame video evaluations. Finally, we integrate our method with the realistic simulator DriveArena to provide more reliable open-loop and closed-loop evaluations for vision-based driving agents. Project Page: https://pjlab-adg.github.io/DriveArena/dreamforge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。