让视频动作迁移更精准,能自动适配内容语义。
MotionAdapter: Video Motion Transfer via Content-Aware Attention Customization
- 通过注意力分析分离动作与外观信息
- 利用DINO模型对齐内容,定制化调整动作
- 支持复杂动作迁移与编辑,如缩放、组合
基于扩散变压器架构的文生视频模型在生成高质量、时序连贯视频方面取得显著进展,但视频间复杂动作迁移仍具挑战。本文提出MotionAdapter,一种内容感知的动作迁移框架,可在DiT-based视频扩散模型中实现鲁棒且语义对齐的动作迁移。核心思想是:1)显式解耦动作与外观;2)根据目标内容自适应定制动作。MotionAdapter首先通过分析3D全注意力模块中的跨帧注意力,提取由注意力生成的动作场,以分离运动信息。为弥合参考视频与目标视频间的语义差距,进一步引入DINO引导的动作定制模块,基于内容对应关系重新排列并优化动作场。定制后的动作场用于指导DiT去噪过程,确保合成视频继承参考动作的同时,保留目标外观和语义。大量实验证明,MotionAdapter在定性与定量评估上均优于现有最优方法。此外,该方法天然支持复杂动作迁移与编辑任务,如缩放、组合等。
原文摘要 · Abstract (English)
Recent advances in diffusion-based text-to-video models, particularly those built on the diffusion transformer architecture, have achieved remarkable progress in generating high-quality and temporally coherent videos. However, transferring complex motions between videos remains challenging. In this work, we present MotionAdapter, a content-aware motion transfer framework that enables robust and semantically aligned motion transfer within DiT-based video diffusion models. Our key insight is that effective motion transfer requires 1) explicit disentanglement of motion from appearance and 2) adaptive customization of motion to target content. MotionAdapter first isolates motion by analyzing cross-frame attention within 3D full-attention modules to extract attention-derived motion fields. To bridge the semantic gap between reference and target videos, we further introduce a DINO-guided motion customization module that rearranges and refines motion fields based on content correspondences. The customized motion field is then used to guide the DiT denoising process, ensuring that the synthesized video inherits the reference motion while preserving target appearance and semantics. Extensive experiments demonstrate that MotionAdapter outperforms state-of-the-art methods in both qualitative and quantitative evaluations. Moreover, MotionAdapter naturely support complex motion transfer and motion editing tasks such as zooming in/out and composition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。