用全局注意力生成多视角驾驶视频,保持时空视点一致
Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention
- 采用4维全局注意力统一建模空间、时间与视角关系
- 在nuScenes上达FVD 37.8,生成视频更真实连续
- 轻量控制器仅占ControlNet 1.1%参数,可精准控制俯视图布局
为自动驾驶训练生成多视角视频受到广泛关注,核心挑战在于跨视角与跨帧的一致性。现有方法通常对空间、时间、视角维度使用解耦注意力机制,难以维持多维度一致性,尤其在快速移动物体出现在不同时间与视角时表现不佳。本文提出CogDriving,一种基于扩散变换器架构的新型网络,引入4维全局注意力模块,实现空间、时间与视角的同步关联。我们还设计了轻量级控制器Micro-Controller,仅需标准ControlNet 1.1%参数,可精确控制鸟瞰图布局。为提升关键物体实例生成质量,提出重加权学习目标,在训练中动态调整物体实例的学习权重。CogDriving在nuScenes验证集上取得FVD 37.8的成绩,证明其生成视频的高度真实性。
原文摘要 · Abstract (English)
Generating multi-view videos for autonomous driving training has recently gained much attention, with the challenge of addressing both cross-view and cross-frame consistency. Existing methods typically apply decoupled attention mechanisms for spatial, temporal, and view dimensions. However, these approaches often struggle to maintain consistency across dimensions, particularly when handling fast-moving objects that appear at different times and viewpoints. In this paper, we present CogDriving, a novel network designed for synthesizing high-quality multi-view driving videos. CogDriving leverages a Diffusion Transformer architecture with holistic-4D attention modules, enabling simultaneous associations across the spatial, temporal, and viewpoint dimensions. We also propose a lightweight controller tailored for CogDriving, i.e., Micro-Controller, which uses only 1.1% of the parameters of the standard ControlNet, enabling precise control over Bird's-Eye-View layouts. To enhance the generation of object instances crucial for autonomous driving, we propose a re-weighted learning objective, dynamically adjusting the learning weights for object instances during training. CogDriving demonstrates strong performance on the nuScenes validation set, achieving an FVD score of 37.8, highlighting its ability to generate realistic driving videos. The project can be found at https://luhannan.github.io/CogDrivingPage/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。