用分割掩码分离多主体运动,实现精准可控的视频动作迁移。
MultiMotion: Multi Subject Video Motion Transfer via Video Diffusion Transformer
- 引入掩码感知注意力流,显式解耦多对象运动特征。
- 在多主体迁移任务中实现语义对齐与时间连贯的动作传递。
- 适用于需要精细控制多个角色动作的视频生成场景。
多物体视频动作迁移对扩散变压器(DiT)架构构成重大挑战,源于固有的运动纠缠与缺乏对象级控制。本文提出MultiMotion,一种新型统一框架以克服这些局限。核心创新为掩码感知注意力运动流(AMF),利用SAM2掩码在DiT流程中显式解耦并控制多个物体的运动特征。此外,我们引入高阶预测-校正求解器RectPC,实现高效且准确的采样,尤其适用于多实体生成。为支持严谨评估,我们构建了首个针对基于DiT的多物体动作迁移的基准数据集。MultiMotion在多个不同物体上实现了精确、语义对齐且时间连贯的动作迁移,同时保持了DiT的高质量与可扩展性。代码见附录。
原文摘要 · Abstract (English)
Multi-object video motion transfer poses significant challenges for Diffusion Transformer (DiT) architectures due to inherent motion entanglement and lack of object-level control. We present MultiMotion, a novel unified framework that overcomes these limitations. Our core innovation is Maskaware Attention Motion Flow (AMF), which utilizes SAM2 masks to explicitly disentangle and control motion features for multiple objects within the DiT pipeline. Furthermore, we introduce RectPC, a high-order predictor-corrector solver for efficient and accurate sampling, particularly beneficial for multi-entity generation. To facilitate rigorous evaluation, we construct the first benchmark dataset specifically for DiT-based multi-object motion transfer. MultiMotion demonstrably achieves precise, semantically aligned, and temporally coherent motion transfer for multiple distinct objects, maintaining DiT's high quality and scalability. The code is in the supp.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。