提出多尺度因果注意力机制,高效生成高分辨率视频。
MSC: Multi-Scale Spatio-Temporal Causal Attention for Autoregressive Video Diffusion
- 引入多尺度空间与时序频率,实现高效注意力计算。
- 利用噪声破坏速率差异,实现因果条件建模,降低计算开销。
- 支持自回归长视频生成,保持帧序自然性,适合视频合成任务。
扩散变换器为视频生成提供了灵活的建模方式。然而,生成具有丰富语义和复杂运动的高分辨率视频仍面临技术挑战且计算成本高昂。与语言类似,视频数据本质上具有自回归特性,因此在模型中使用具有双向依赖关系的注意力机制是反直觉的。为此,本文提出多尺度因果(MSC)框架。具体而言,我们在空间维度引入多分辨率,在时间维度引入高低频分层,以实现高效的注意力计算。此外,通过受控方式组合多尺度注意力模块,使扩散训练能够基于噪声图像帧进行因果条件建模,其原理在于不同分辨率上的噪声以不同速率破坏信息。我们从理论上证明该方法可显著降低计算复杂度并提升训练效率。所提出的因果注意力扩散框架还可用于自回归长视频生成,不违背帧序列的自然顺序。
原文摘要 · Abstract (English)
Diffusion transformers enable flexible generative modeling for video. However, it is still technically challenging and computationally expensive to generate high-resolution videos with rich semantics and complex motion. Similar to languages, video data are also auto-regressive by nature, so it is counter-intuitive to use attention mechanism with bi-directional dependency in the model. Here we propose a Multi-Scale Causal (MSC) framework to address these problems. Specifically, we introduce multiple resolutions in the spatial dimension and high-low frequencies in the temporal dimension to realize efficient attention calculation. Furthermore, attention blocks on multiple scales are combined in a controlled way to allow causal conditioning on noisy image frames for diffusion training, based on the idea that noise destroys information at different rates on different resolutions. We theoretically show that our approach can greatly reduce the computational complexity and enhance the efficiency of training. The causal attention diffusion framework can also be used for auto-regressive long video generation, without violating the natural order of frame sequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。