arXiv:2603.01000cs.CV2026-03被引 2

让多物体在视频中各自按不同动作动起来,支持自由组合

Let Your Image Move with Your Motion! -- Implicit Multi-Object Multi-Motion Transfer

  • 用对象专属掩码解耦运动注意力,避免动作混淆
  • 在多个参考视频上实现精确的多物体动作迁移,效果领先
  • 适合需要灵活控制多个角色动作的视频生成场景

运动迁移已成为可控视频生成的重要方向,但现有方法多集中于单物体场景,在多个物体需不同运动模式时表现不佳。本文提出 FlexiMMT,首个隐式图像到视频(I2V)多物体多运动迁移框架,可独立提取运动表征并精准分配至不同物体,支持灵活重组与任意运动-物体映射。为解决跨物体运动纠缠问题,引入运动解耦掩码注意力机制,利用对象特定掩码约束注意力,确保运动与文本标记仅影响指定区域。进一步提出差异化掩码传播机制,直接从扩散注意力中推导出对象特定掩码,并高效跨帧传播。大量实验表明,FlexiMMT 在基于 I2V 的多物体多运动迁移任务中实现了精确、组合性强且达到顶尖水平的表现。

原文摘要 · Abstract (English)

Motion transfer has emerged as a promising direction for controllable video generation, yet existing methods largely focus on single-object scenarios and struggle when multiple objects require distinct motion patterns. In this work, we present FlexiMMT, the first implicit image-to-video (I2V) motion transfer framework that explicitly enables multi-object, multi-motion transfer. Given a static multi-object image and multiple reference videos, FlexiMMT independently extracts motion representations and accurately assigns them to different objects, supporting flexible recombination and arbitrary motion-to-object mappings. To address the core challenge of cross-object motion entanglement, we introduce a Motion Decoupled Mask Attention Mechanism that uses object-specific masks to constrain attention, ensuring that motion and text tokens only influence their designated regions. We further propose a Differentiated Mask Propagation Mechanism that derives object-specific masks directly from diffusion attention and progressively propagates them across frames efficiently. Extensive experiments demonstrate that FlexiMMT achieves precise, compositional, and state-of-the-art performance in I2V-based multi-object multi-motion transfer. Our project page is: https://ethan-li123.github.io/FlexiMMT_page/

视频生成运动迁移多物体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。