arXiv:2504.01724cs.CVcs.AI2025-04ICCV被引 41

用混合引导提升人像动画的表达力与稳定性

DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidance

论文配图:DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidance
图 1 · 摘自论文原文
  • 融合面部隐式表征、3D头身与骨架,实现精准动作控制
  • 多尺度渐进训练,适配从半身到全身的多种视角
  • 结合时序运动模式与视觉参考,保持长时间动作连贯性

现有基于图像的人体动画方法虽能生成逼真的身体与面部动作,但在细粒度整体可控性、多尺度适应性及长期时间一致性方面仍存在显著缺陷,导致表达力和鲁棒性不足。本文提出基于扩散变换器(DiT)的DreamActor-M1框架,采用混合引导机制克服上述问题。在动作引导方面,融合隐式面部表征、3D头球与3D身体骨骼,实现对表情与动作的鲁棒控制,并生成富有表现力且身份一致的动画。在尺度自适应方面,通过使用不同分辨率与尺度的数据进行渐进式训练,支持从半身照到全身像的多场景输入。在外观引导方面,整合连续帧中的运动模式与互补视觉参考,确保复杂动作中未见区域的长期时序一致性。实验表明,该方法在半身、上半身及全身生成任务中均优于当前最优模型,展现出卓越的表现力与长期稳定性。

原文摘要 · Abstract (English)

While recent image-based human animation methods achieve realistic body and facial motion synthesis, critical gaps remain in fine-grained holistic controllability, multi-scale adaptability, and long-term temporal coherence, which leads to their lower expressiveness and robustness. We propose a diffusion transformer (DiT) based framework, DreamActor-M1, with hybrid guidance to overcome these limitations. For motion guidance, our hybrid control signals that integrate implicit facial representations, 3D head spheres, and 3D body skeletons achieve robust control of facial expressions and body movements, while producing expressive and identity-preserving animations. For scale adaptation, to handle various body poses and image scales ranging from portraits to full-body views, we employ a progressive training strategy using data with varying resolutions and scales. For appearance guidance, we integrate motion patterns from sequential frames with complementary visual references, ensuring long-term temporal coherence for unseen regions during complex movements. Experiments demonstrate that our method outperforms the state-of-the-art works, delivering expressive results for portraits, upper-body, and full-body generation with robust long-term consistency. Project Page: https://grisoon.github.io/DreamActor-M1/.

人像动画扩散模型动作控制时序一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。