arXiv:2412.05275cs.CVcs.AI2024-12被引 17

无需训练,用注意力机制实现视频运动精准迁移

MotionFlow: Attention-Driven Motion Transfer in Video Diffusion Models

  • 通过交叉注意力图捕捉时空动态,实现运动信息精准转移
  • 在剧烈场景变化下仍保持高保真与运动一致性,超越现有方法
  • 即插即用,适用于各类预训练视频扩散模型

文本生成视频模型已展现出生成多样且吸引人视频内容的强大能力,标志着生成式AI的重要进展。然而,这些模型普遍缺乏对运动模式的细粒度控制,限制了实际应用。我们提出MotionFlow,一种用于视频扩散模型中运动迁移的新框架。该方法利用交叉注意力图精确捕捉并操纵空间与时间动态,实现跨多种场景的无缝运动转移。本方法无需训练,仅在测试时借助预训练视频扩散模型的内在能力即可工作。相比传统方法在复杂场景变换中难以兼顾整体变化与运动一致性,MotionFlow通过其基于注意力的机制成功应对此类挑战。定性和定量实验表明,即便在剧烈场景改变下,MotionFlow在保真度和多样性方面均显著优于现有模型。

原文摘要 · Abstract (English)

Text-to-video models have demonstrated impressive capabilities in producing diverse and captivating video content, showcasing a notable advancement in generative AI. However, these models generally lack fine-grained control over motion patterns, limiting their practical applicability. We introduce MotionFlow, a novel framework designed for motion transfer in video diffusion models. Our method utilizes cross-attention maps to accurately capture and manipulate spatial and temporal dynamics, enabling seamless motion transfers across various contexts. Our approach does not require training and works on test-time by leveraging the inherent capabilities of pre-trained video diffusion models. In contrast to traditional approaches, which struggle with comprehensive scene changes while maintaining consistent motion, MotionFlow successfully handles such complex transformations through its attention-based mechanism. Our qualitative and quantitative experiments demonstrate that MotionFlow significantly outperforms existing models in both fidelity and versatility even during drastic scene alterations.

视频生成扩散模型运动迁移注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。