让视频生成持续学习新概念,不遗忘旧内容。
Bring Your Dreams to Life: Continual Text-to-Video Customization

- 设计新模型实现视频生成的持续学习。
- 在多个基准上优于现有方法,避免遗忘和忽略新概念。
- 适合需要长期个性化视频生成的用户。
定制化文本到视频生成(CTVG)近期在根据用户特定文本生成定制视频方面取得显著进展。然而,大多数CTVG方法假设个性化概念保持静态,无法随时间增量扩展。此外,它们在持续学习新概念(包括主体和动作)时,容易出现遗忘和概念忽略问题。为解决上述挑战,我们提出一种新型持续定制视频扩散模型(CCVD),可通过应对遗忘和概念忽略,持续学习新概念以生成跨多种任务的视频。为缓解灾难性遗忘,引入概念特定属性保留模块和任务感知概念聚合策略,可在训练中捕捉旧概念的独特特征与身份,并在测试时根据相关性整合所有旧概念的主体与动作适配器。此外,为缓解概念忽略,开发可控条件合成机制,通过层特定区域注意力引导的噪声估计增强局部特征,并对齐视频上下文与用户条件。大量实验表明,我们的CCVD在DreamVideo和Wan 2.1两个骨干网络上均优于现有基线。代码已公开于https://github.com/JiahuaDong/CCVD。
原文摘要 · Abstract (English)
Customized text-to-video generation (CTVG) has recently witnessed great progress in generating tailored videos from user-specific text. However, most CTVG methods assume that personalized concepts remain static and do not expand incrementally over time. Additionally, they struggle with forgetting and concept neglect when continuously learning new concepts, including subjects and motions. To resolve the above challenges, we develop a novel Continual Customized Video Diffusion (CCVD) model, which can continuously learn new concepts to generate videos across various text-to-video generation tasks by tackling forgetting and concept neglect. To address catastrophic forgetting, we introduce a concept-specific attribute retention module and a task-aware concept aggregation strategy. They can capture the unique characteristics and identities of old concepts during training, while combining all subject and motion adapters of old concepts based on their relevance during testing. Besides, to tackle concept neglect, we develop a controllable conditional synthesis to enhance regional features and align video contexts with user conditions, by incorporating layer-specific region attention-guided noise estimation. Extensive experimental comparisons demonstrate that our CCVD outperforms existing CTVG baselines on both the DreamVideo and Wan 2.1 backbones. The code is available at https://github.com/JiahuaDong/CCVD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。