arXiv:2411.18281cs.CV2024-11AAAI

让视频人物动作可精确调节,还能保持脸型不变。

MotionCharacter: Fine-Grained Motion Controllable Human Video Generation

  • 把动作类型和强度分开控制,用文字和光流数据调节
  • 生成视频身份一致,动作精准匹配指定强度
  • 适合虚拟偶像、微表情动画等高精度场景

个性化文本到视频生成虽有进展,但难以精细控制动作强度。问题源于动作语义与幅度在粗略文本中纠缠,限制了细微动作合成,影响虚拟角色或微表情应用。为此,我们提出MotionCharacter框架,将运动解耦为动作类型与强度两个独立可控成分。核心创新:(1) 运动控制模块利用文本指定动作类型,通过光流导出的量化指标调节强度,并采用区域感知损失聚焦运动区域;(2) 身份内容插入模块配合身份一致性损失,保障动态中身份稳定。为支持训练,我们构建了包含动作与面部特征详细标注的大规模数据集Human-Motion。大量实验表明,MotionCharacter在身份一致性与动作类型/强度精确性上显著优于现有方法。

原文摘要 · Abstract (English)

Recent advancements in personalized Text-to-Video (T2V) generation have made significant strides in synthesizing character-specific content. However, these methods face a critical limitation: the inability to perform fine-grained control over motion intensity. This limitation stems from an inherent entanglement of action semantics and their corresponding magnitudes within coarse textual descriptions, hindering the generation of nuanced human videos and limiting their applicability in scenarios demanding high precision, such as animating virtual avatars or synthesizing subtle micro-expressions. Furthermore, existing approaches often struggle to preserve high identity fidelity when other attributes are modified. To address these challenges, we introduce MotionCharacter, a framework for high-fidelity human video generation with precise motion control. At its core, MotionCharacter explicitly decouples motion into two independently controllable components: action type and motion intensity. This is achieved through two key technical contributions: (1) a Motion Control Module that leverages textual phrases to specify the action type and a quantifiable metric derived from optical flow to modulate its intensity, guided by a region-aware loss that localizes motion to relevant subject areas; and (2) an ID Content Insertion Module coupled with an ID-Consistency loss to ensure robust identity preservation during dynamic motions. To facilitate training for such fine-grained control, we also curate Human-Motion, a new large-scale dataset with detailed annotations for both motion and facial features. Extensive experiments demonstrate that MotionCharacter achieves substantial improvements over existing methods. Our framework excels in generating videos that are not only identity-consistent but also precisely adhere to specified motion types and intensities.

视频生成动作控制身份保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。