用单段视频让视频生成模型学会动态角色的外观与动作
Dynamic Concepts Personalization from Single Videos
- 先用无序帧学外观,再加运动残差学动作
- 在扩散模型中实现外观与动作的联合建模
- 适合需要个性化动态角色生成的研究者
个性化生成文本到图像模型已取得显著进展,但将个性化扩展到文本到视频模型面临独特挑战。与静态概念不同,文本到视频模型的个性化可捕捉动态概念,即不仅由外观定义,还包含运动特征。本文提出Set-and-Sequence框架,用于基于扩散变换器(DiTs)的生成视频模型的动态概念个性化。该方法在不显式分离空间与时间特征的架构中引入时空权重空间。第一阶段,使用视频中无序帧微调低秩适应(LoRA)层,学习一个不受时间干扰的身份LoRA基,表示外观;第二阶段,在冻结身份LoRA的基础上,通过运动残差增强其系数,并在完整视频序列上微调,以捕捉运动动态。Set-and-Sequence框架在视频模型输出域中有效嵌入动态概念,实现前所未有的可编辑性和组合性,为动态概念个性化设定新基准。
原文摘要 · Abstract (English)
Personalizing generative text-to-image models has seen remarkable progress, but extending this personalization to text-to-video models presents unique challenges. Unlike static concepts, personalizing text-to-video models has the potential to capture dynamic concepts, i.e., entities defined not only by their appearance but also by their motion. In this paper, we introduce Set-and-Sequence, a novel framework for personalizing Diffusion Transformers (DiTs)-based generative video models with dynamic concepts. Our approach imposes a spatio-temporal weight space within an architecture that does not explicitly separate spatial and temporal features. This is achieved in two key stages. First, we fine-tune Low-Rank Adaptation (LoRA) layers using an unordered set of frames from the video to learn an identity LoRA basis that represents the appearance, free from temporal interference. In the second stage, with the identity LoRAs frozen, we augment their coefficients with Motion Residuals and fine-tune them on the full video sequence, capturing motion dynamics. Our Set-and-Sequence framework results in a spatio-temporal weight space that effectively embeds dynamic concepts into the video model's output domain, enabling unprecedented editability and compositionality while setting a new benchmark for personalizing dynamic concepts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。