arXiv:2410.13720cs.CVcs.AI2024-10被引 633

Movie Gen 能生成16秒高清视频并同步音频,支持个性化定制和精准编辑。

Movie Gen: A Cast of Media Foundation Models

论文配图:Movie Gen: A Cast of Media Foundation Models
图 1 · 摘自论文原文
  • 基于300亿参数模型,支持1080p高清视频生成与多比例适配。
  • 在文本到视频等五项任务中达到当前最佳性能,最长生成16秒视频。
  • 适用于视频创作、影视制作及个性化内容生成的科研与工业场景。

我们提出 Movie Gen,一系列可生成高质量1080p高清视频(支持多种画幅比)并同步音频的媒体基础模型。此外,模型还具备基于指令的精确视频编辑能力,以及根据用户图像生成个性化视频的功能。在文本到视频合成、视频个性化、视频编辑、视频到音频生成和文本到音频生成等多个任务上,模型均达到新基准。最大视频生成模型为30B参数的Transformer,最大上下文长度达73K视频标记,对应16秒时长(16帧/秒)。我们通过架构创新、潜在空间设计、训练目标优化、数据筛选、评估协议、并行化技术及推理加速等多项改进,有效利用大规模预训练数据、模型规模与训练算力。相关视频已公开于 https://go.fb.me/MovieGenResearchVideos。

原文摘要 · Abstract (English)

We present Movie Gen, a cast of foundation models that generates high-quality, 1080p HD videos with different aspect ratios and synchronized audio. We also show additional capabilities such as precise instruction-based video editing and generation of personalized videos based on a user's image. Our models set a new state-of-the-art on multiple tasks: text-to-video synthesis, video personalization, video editing, video-to-audio generation, and text-to-audio generation. Our largest video generation model is a 30B parameter transformer trained with a maximum context length of 73K video tokens, corresponding to a generated video of 16 seconds at 16 frames-per-second. We show multiple technical innovations and simplifications on the architecture, latent spaces, training objectives and recipes, data curation, evaluation protocols, parallelization techniques, and inference optimizations that allow us to reap the benefits of scaling pre-training data, model size, and training compute for training large scale media generation models. We hope this paper helps the research community to accelerate progress and innovation in media generation models. All videos from this paper are available at https://go.fb.me/MovieGenResearchVideos.

视频生成扩散模型个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。