用50亿参数的混合模型突破状态空间模型在图像视频生成中的极限
Pushing the Boundaries of State Space Models for Image and Video Generation
- 构建双向高效混合架构,融合SSM与自注意力机制
- 实现2K分辨率图像和360p/8秒动态视频生成
- 适合追求长序列生成效率的视觉生成研究者
尽管Transformer已成为视觉生成的主流架构,线性注意力模型如状态空间模型(SSM)因其处理长视觉序列的高效性日益受到关注。然而,其依赖有限递归状态并强制因果性,难以一致建模多维视觉数据,限制了其在生成非因果长序列上的能力。本文通过构建迄今最大的扩散SSM-Transformer混合模型(50亿参数),基于亚二次双向Hydra与自注意力机制,成功生成最高达2K分辨率的图像及360p、8秒(16 FPS)的视频。结果表明,模型能忠实响应复杂文本提示,生成时间上一致且高动态的视频,展现了SSM在视觉生成任务中的巨大潜力。
原文摘要 · Abstract (English)
While Transformers have become the dominant architecture for visual generation, linear attention models, such as the state-space models (SSM), are increasingly recognized for their efficiency in processing long visual sequences. However, the essential efficiency of these models comes from formulating a limited recurrent state, enforcing causality among tokens that are prone to inconsistent modeling of N-dimensional visual data, leaving questions on their capacity to generate long non-causal sequences. In this paper, we explore the boundary of SSM on image and video generation by building the largest-scale diffusion SSM-Transformer hybrid model to date (5B parameters) based on the sub-quadratic bi-directional Hydra and self-attention, and generate up to 2K images and 360p 8 seconds (16 FPS) videos. Our results demonstrate that the model can produce faithful results aligned with complex text prompts and temporal consistent videos with high dynamics, suggesting the great potential of using SSMs for visual generation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。