arXiv:2412.15646cs.CV2024-12中稿 · AAAI被引 14

让视频生成同时定制动作与外观,解决多参考融合的伪影问题。

CustomTTT: Motion and Appearance Customized Video Generation via Test-Time Training

  • 分离控制动作与外观的LoRA层,精准定位定制位置。
  • 测试时训练动态更新参数,融合多参考内容无明显伪影。
  • 适合需要个性化角色和动作生成的创意应用。

得益于大规模文本-视频对的预训练,当前的文本到视频(T2V)扩散模型能根据文本描述生成高质量视频。此外,通过参数高效的微调方法LoRA,给定参考图像或视频后可生成特定主题或动作的定制化内容。然而,将多个不同参考所训练的定制概念合并到单一网络中时,会出现明显伪影。为此,我们提出CustomTTT,可轻松联合定制给定视频的动作与外观。具体而言,我们分析了提示词在现有视频扩散模型中的影响,发现仅需在特定层应用LoRA即可完成外观与动作的定制。由于每个LoRA独立训练,我们提出一种新颖的测试时训练技术,在组合后利用已训练的定制模型动态更新参数。通过详尽实验验证了方法的有效性,我们的方法在定性和定量评估中均优于多个最先进工作。

原文摘要 · Abstract (English)

Benefiting from large-scale pre-training of text-video pairs, current text-to-video (T2V) diffusion models can generate high-quality videos from the text description. Besides, given some reference images or videos, the parameter-efficient fine-tuning method, i.e. LoRA, can generate high-quality customized concepts, e.g., the specific subject or the motions from a reference video. However, combining the trained multiple concepts from different references into a single network shows obvious artifacts. To this end, we propose CustomTTT, where we can joint custom the appearance and the motion of the given video easily. In detail, we first analyze the prompt influence in the current video diffusion model and find the LoRAs are only needed for the specific layers for appearance and motion customization. Besides, since each LoRA is trained individually, we propose a novel test-time training technique to update parameters after combination utilizing the trained customized models. We conduct detailed experiments to verify the effectiveness of the proposed methods. Our method outperforms several state-of-the-art works in both qualitative and quantitative evaluations.

视频生成扩散模型个性化定制测试时训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。