90亿参数模型,用文本到图像预训练模型改进视频生成效果
CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

- 继承文本到图像模型权重,降低视频生成训练成本
- 采用多帧率分层训练,提升文本与视频语义对齐
- 首个开源大模型,在机器与人工评估中全面领先
大规模预训练变换器在文本(GPT-3)和文本到图像(DALL-E、CogView)生成中取得突破,但在视频生成领域仍面临挑战:高昂的计算成本使从零训练不可行;文本-视频数据集稀缺且相关性弱,难以理解复杂运动语义。本文提出90亿参数的Transformer模型CogVideo,通过继承预训练的文本到图像模型CogView2进行训练,并提出多帧率分层训练策略以更好对齐文本与视频片段。作为(可能)首个开源的大规模预训练文本到视频生成模型,CogVideo在机器与人工评估中均大幅超越现有公开模型。
原文摘要 · Abstract (English)
Large-scale pretrained transformers have created milestones in text (GPT-3) and text-to-image (DALL-E and CogView) generation. Its application to video generation is still facing many challenges: The potential huge computation cost makes the training from scratch unaffordable; The scarcity and weak relevance of text-video datasets hinder the model understanding complex movement semantics. In this work, we present 9B-parameter transformer CogVideo, trained by inheriting a pretrained text-to-image model, CogView2. We also propose multi-frame-rate hierarchical training strategy to better align text and video clips. As (probably) the first open-source large-scale pretrained text-to-video model, CogVideo outperforms all publicly available models at a large margin in machine and human evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。