基于DiT的多视角视频生成框架,实现高一致性控制。
DiVE: DiT-based Video Generation with Enhanced Control
- 用无参空间视图膨胀注意力机制保证多视角一致性
- 在nuScenes数据集上实现长时序、精准控制的视频生成
- 适合自动驾驶场景下复杂边缘案例的视频合成
在自动驾驶场景中生成高质量、时间一致的视频仍面临挑战,尤其是在边缘案例中出现异常操作。尽管已有基于扩散变换器(DiT)的视频生成方法被提出,但针对多视角视频生成场景的研究仍属空白。本文首次提出一个专为生成时空与多视角一致视频设计的DiT框架,可精确匹配给定的鸟瞰图布局控制。该框架采用无参的空间视图膨胀注意力机制,确保跨视角一致性,并集成联合跨注意力模块与ControlNet-Transformer以提升控制精度。我们在nuScenes数据集上进行了广泛定性对比,尤其聚焦于最具挑战性的边缘案例。结果表明,所提方法在复杂条件下能有效生成长时序、可控且高度一致的视频。
原文摘要 · Abstract (English)
Generating high-fidelity, temporally consistent videos in autonomous driving scenarios faces a significant challenge, e.g. problematic maneuvers in corner cases. Despite recent video generation works are proposed to tackcle the mentioned problem, i.e. models built on top of Diffusion Transformers (DiT), works are still missing which are targeted on exploring the potential for multi-view videos generation scenarios. Noticeably, we propose the first DiT-based framework specifically designed for generating temporally and multi-view consistent videos which precisely match the given bird's-eye view layouts control. Specifically, the proposed framework leverages a parameter-free spatial view-inflated attention mechanism to guarantee the cross-view consistency, where joint cross-attention modules and ControlNet-Transformer are integrated to further improve the precision of control. To demonstrate our advantages, we extensively investigate the qualitative comparisons on nuScenes dataset, particularly in some most challenging corner cases. In summary, the effectiveness of our proposed method in producing long, controllable, and highly consistent videos under difficult conditions is proven to be effective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。