arXiv:2205.15868cs.CVcs.CL2022-05ICLR被引 1.2k

90亿参数模型,用文本到图像预训练模型改进视频生成效果

CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

论文配图:CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
图 1 · 摘自论文原文
  • 继承文本到图像模型权重,降低视频生成训练成本
  • 采用多帧率分层训练,提升文本与视频语义对齐
  • 首个开源大模型,在机器与人工评估中全面领先

大规模预训练变换器在文本(GPT-3)和文本到图像(DALL-E、CogView)生成中取得突破,但在视频生成领域仍面临挑战:高昂的计算成本使从零训练不可行;文本-视频数据集稀缺且相关性弱,难以理解复杂运动语义。本文提出90亿参数的Transformer模型CogVideo,通过继承预训练的文本到图像模型CogView2进行训练,并提出多帧率分层训练策略以更好对齐文本与视频片段。作为(可能)首个开源的大规模预训练文本到视频生成模型,CogVideo在机器与人工评估中均大幅超越现有公开模型。

原文摘要 · Abstract (English)

Large-scale pretrained transformers have created milestones in text (GPT-3) and text-to-image (DALL-E and CogView) generation. Its application to video generation is still facing many challenges: The potential huge computation cost makes the training from scratch unaffordable; The scarcity and weak relevance of text-video datasets hinder the model understanding complex movement semantics. In this work, we present 9B-parameter transformer CogVideo, trained by inheriting a pretrained text-to-image model, CogView2. We also propose multi-frame-rate hierarchical training strategy to better align text and video clips. As (probably) the first open-source large-scale pretrained text-to-video model, CogVideo outperforms all publicly available models at a large margin in machine and human evaluations.

视频生成Transformer大模型文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。