arXiv:2412.07776cs.CVcs.AI2024-12CVPR被引 42

用注意力流实现扩散模型的零样本视频动作迁移。

Video Motion Transfer with Diffusion Transformers

  • 基于预训练扩散变换器提取跨帧注意力流,生成运动信号。
  • 无需训练,通过优化损失函数实现动作复现,性能超越现有方法。
  • 适用于零样本动作迁移,适合视频生成与动画创作研究者。

我们提出DiTFlow,一种将参考视频动作迁移到新合成视频中的方法,专为扩散变换器(DiT)设计。首先利用预训练的DiT处理参考视频,分析跨帧注意力图并提取一种称为注意力运动流(AMF)的局部运动信号。随后,在优化基础上、无须训练地引导潜空间去噪过程,通过最小化我们的AMF损失来生成复现参考动作的视频。此外,我们将该优化策略应用于变换器的位置嵌入,显著提升零样本动作迁移能力。在多个指标和人工评估中,DiTFlow均优于近期发表的方法。

原文摘要 · Abstract (English)

We propose DiTFlow, a method for transferring the motion of a reference video to a newly synthesized one, designed specifically for Diffusion Transformers (DiT). We first process the reference video with a pre-trained DiT to analyze cross-frame attention maps and extract a patch-wise motion signal called the Attention Motion Flow (AMF). We guide the latent denoising process in an optimization-based, training-free, manner by optimizing latents with our AMF loss to generate videos reproducing the motion of the reference one. We also apply our optimization strategy to transformer positional embeddings, granting us a boost in zero-shot motion transfer capabilities. We evaluate DiTFlow against recently published methods, outperforming all across multiple metrics and human evaluation.

视频生成扩散模型动作迁移Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。