arXiv:2209.14792cs.CVcs.AI2022-09ICLR被引 2.2k

无需图文数据,直接将文生图能力扩展到文生视频。

Make-A-Video: Text-to-Video Generation without Text-Video Data

论文配图:Make-A-Video: Text-to-Video Generation without Text-Video Data
图 1 · 摘自论文原文
  • 复用已有图文模型,用视频无监督学习运动规律。
  • 生成视频在分辨率、帧率和文本忠实度上均达新高。
  • 适合需要高质量视频生成的创作者与研究者。

我们提出 Make-A-Video——一种将文本到图像(T2I)生成的最新进展直接迁移到文本到视频(T2V)生成的方法。核心思路是:从配对图文数据中学习世界外观及描述方式,从无监督视频中学习世界运动规律。该方法有三大优势:(1) 加速 T2V 模型训练(无需从头学习视觉与多模态表示),(2) 不依赖成对图文数据,(3) 生成视频继承当前图像生成模型的多样性与想象力。我们设计了一种简单而有效的空间-时间模块,在现有 T2I 模型基础上构建新架构:首先分解并近似全时序 U-Net 和注意力张量,实现时空分离;其次构建空间-时间流水线,通过视频解码器、插值模型与两个超分模型生成高分辨率、高帧率视频,支持多种下游应用。在空间与时间分辨率、文本忠实度、生成质量等方面,Make-A-Video 均在定性与定量评估中达到新基准。

原文摘要 · Abstract (English)

We propose Make-A-Video -- an approach for directly translating the tremendous recent progress in Text-to-Image (T2I) generation to Text-to-Video (T2V). Our intuition is simple: learn what the world looks like and how it is described from paired text-image data, and learn how the world moves from unsupervised video footage. Make-A-Video has three advantages: (1) it accelerates training of the T2V model (it does not need to learn visual and multimodal representations from scratch), (2) it does not require paired text-video data, and (3) the generated videos inherit the vastness (diversity in aesthetic, fantastical depictions, etc.) of today's image generation models. We design a simple yet effective way to build on T2I models with novel and effective spatial-temporal modules. First, we decompose the full temporal U-Net and attention tensors and approximate them in space and time. Second, we design a spatial temporal pipeline to generate high resolution and frame rate videos with a video decoder, interpolation model and two super resolution models that can enable various applications besides T2V. In all aspects, spatial and temporal resolution, faithfulness to text, and quality, Make-A-Video sets the new state-of-the-art in text-to-video generation, as determined by both qualitative and quantitative measures.

文生视频扩散模型无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。