arXiv:2412.18966cs.CVcs.AI2024-12被引 4

让视频生成模型像人一样持续学习,逐步变强。

ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement

  • 通过扩展模型规模和引入大语言模型增强语义理解
  • 在复杂提示下生成效果显著提升,实现更好语义对齐
  • 适合资源有限但需持续优化视频生成的团队使用

文本到视频(T2V)生成近年来备受关注,但从零训练模型成本高昂,且在计算资源受限时性能仍有提升空间。本文探索了文本到视频模型的持续通用预训练,使模型能基于已有基础不断“成长”,类似人类基于过往经验学习新知识。当前对T2V持续预训练的研究尚不充分。为此,我们提出ModelGrow,系统性地将该任务分解为两大关键方面:提升模型容量与增强语义理解。针对模型容量,我们引入多种新方法扩展模型规模,以存储新知识并提升生成质量;针对语义理解,我们提出利用大语言模型作为高级文本编码器,集成至T2V模型中,增强语言解析能力,使生成结果更贴合详细提示。该方法显著提升了复杂提示下的语义对齐效果。大量实验验证了其有效性。代码与模型将公开发布。

原文摘要 · Abstract (English)

Text-to-video (T2V) generation has gained significant attention recently. However, the costs of training a T2V model from scratch remain persistently high, and there is considerable room for improving the generation performance, especially under limited computation resources. This work explores the continual general pre-training of text-to-video models, enabling the model to "grow" its abilities based on a pre-trained foundation, analogous to how humans acquire new knowledge based on past experiences. There is a lack of extensive study of the continual pre-training techniques in T2V generation. In this work, we take the initial step toward exploring this task systematically and propose ModelGrow. Specifically, we break this task into two key aspects: increasing model capacity and improving semantic understanding. For model capacity, we introduce several novel techniques to expand the model size, enabling it to store new knowledge and improve generation performance. For semantic understanding, we propose a method that leverages large language models as advanced text encoders, integrating them into T2V models to enhance language comprehension and guide generation results according to detailed prompts. This approach enables the model to achieve better semantic alignment, particularly in response to complex user prompts. Extensive experiments demonstrate the effectiveness of our method across various metrics. The source code and the model of ModelGrow will be publicly available.

视频生成持续学习大语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。