无需训练即可实现视频扩散模型中的零样本动作迁移。
MotionShop: Zero-Shot Motion Transfer in Video Diffusion Models with Mixture of Score Guidance
- 通过分数引导混合框架分离动作与内容信息,实现动作迁移。
- 在200个源视频上成功完成单/多物体及复杂镜头运动迁移。
- 适合需要快速动作重用的视频生成研究者使用。
本文提出首个基于扩散变换器的零样本动作迁移方法——分数引导混合(MSG),该方法从理论上重构条件分数,将扩散模型中的动作分数与内容分数解耦。通过将动作迁移建模为势能混合,MSG 在不改变场景结构的前提下,自然保留原始画面布局,并支持创造性场景变换,同时维持动作模式的完整性。该采样方法可直接作用于预训练视频扩散模型,无需额外训练或微调。大量实验表明,MSG 能有效处理单物体、多物体及跨物体动作迁移,以及复杂镜头运动迁移。此外,我们构建了首个动作迁移数据集 MotionBench,包含200个源视频和1000个迁移动作,覆盖单/多物体转移与复杂相机运动。
原文摘要 · Abstract (English)
In this work, we propose the first motion transfer approach in diffusion transformer through Mixture of Score Guidance (MSG), a theoretically-grounded framework for motion transfer in diffusion models. Our key theoretical contribution lies in reformulating conditional score to decompose motion score and content score in diffusion models. By formulating motion transfer as a mixture of potential energies, MSG naturally preserves scene composition and enables creative scene transformations while maintaining the integrity of transferred motion patterns. This novel sampling operates directly on pre-trained video diffusion models without additional training or fine-tuning. Through extensive experiments, MSG demonstrates successful handling of diverse scenarios including single object, multiple objects, and cross-object motion transfer as well as complex camera motion transfer. Additionally, we introduce MotionBench, the first motion transfer dataset consisting of 200 source videos and 1000 transferred motions, covering single/multi-object transfers, and complex camera motions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。