arXiv:2504.02451cs.CV2025-04CVPR被引 14

ConMo实现零样本运动解耦重组,精准迁移多主体视频动作。

ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer

  • 仅用主体掩码分离视频中人物与背景运动,实现解耦控制。
  • 在不同体型主体上保持动作保真度与语义一致性,显著优于现有方法。
  • 支持主体大小位置编辑、移除及镜头运动模拟,适用场景广泛。

文本到视频生成的发展使动作迁移成为可能,能基于已有画面控制视频动作。然而当前方法存在两大局限:1)难以处理多主体视频,无法精准迁移特定主体动作;2)在迁移至不同形状主体时,动作多样性与准确性下降。为此,我们提出 extbf{ConMo},一种零样本框架,可解耦并重组主体与相机运动。ConMo仅通过主体掩码,从源视频复杂轨迹中分离个体主体与背景运动信号,并重新组装以生成目标视频。该方法提升了跨不同主体的动作控制精度,增强了多主体场景表现力。此外,我们在重组阶段引入软引导机制,调节原始动作保留程度以适应主体形状约束,促进形态适配与语义转换。相较于以往方法,ConMo拓展了丰富应用,包括主体尺寸/位置编辑、主体移除、语义修改和相机运动模拟。大量实验表明,ConMo在动作保真度与语义一致性方面显著超越现有最优方法。代码已公开于https://github.com/Andyplus1/ConMo。

原文摘要 · Abstract (English)

The development of Text-to-Video (T2V) generation has made motion transfer possible, enabling the control of video motion based on existing footage. However, current methods have two limitations: 1) struggle to handle multi-subjects videos, failing to transfer specific subject motion; 2) struggle to preserve the diversity and accuracy of motion as transferring to subjects with varying shapes. To overcome these, we introduce \textbf{ConMo}, a zero-shot framework that disentangle and recompose the motions of subjects and camera movements. ConMo isolates individual subject and background motion cues from complex trajectories in source videos using only subject masks, and reassembles them for target video generation. This approach enables more accurate motion control across diverse subjects and improves performance in multi-subject scenarios. Additionally, we propose soft guidance in the recomposition stage which controls the retention of original motion to adjust shape constraints, aiding subject shape adaptation and semantic transformation. Unlike previous methods, ConMo unlocks a wide range of applications, including subject size and position editing, subject removal, semantic modifications, and camera motion simulation. Extensive experiments demonstrate that ConMo significantly outperforms state-of-the-art methods in motion fidelity and semantic consistency. The code is available at https://github.com/Andyplus1/ConMo.

动作迁移视频生成解耦控制零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。