arXiv:2503.18860cs.CV2025-03CVPR被引 62

用隐式控制让单图人像动起来,表情动作自然连贯。

HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation

  • 用隐式表示编码动态信息,解耦身份与动作
  • 生成视频时保持动作连贯、细节丰富
  • 支持多种画风,适合影视动画和虚拟角色

我们提出 HunyuanPortrait,一种基于扩散模型的条件控制方法,采用隐式表示实现高度可控且逼真的肖像动画。给定一张参考肖像图和视频驱动模板,该方法可依据视频中的人物表情与头部姿态,驱动参考图像中的人物进行动画生成。框架利用预训练编码器实现视频中运动信息与身份特征的解耦。通过隐式表示编码运动信息,并在动画阶段作为控制信号注入。基于稳定视频扩散模型(Stable Video Diffusion)构建主干网络,设计适配层,通过注意力机制将控制信号注入去噪UNet,提升空间细节丰富度与时间一致性。HunyuanPortrait展现出强泛化能力,可在不同图像风格下有效分离外观与运动。实验表明,其在时间一致性与可控性上优于现有方法。项目地址:https://kkakkkka.github.io/HunyuanPortrait。

原文摘要 · Abstract (English)

We introduce HunyuanPortrait, a diffusion-based condition control method that employs implicit representations for highly controllable and lifelike portrait animation. Given a single portrait image as an appearance reference and video clips as driving templates, HunyuanPortrait can animate the character in the reference image by the facial expression and head pose of the driving videos. In our framework, we utilize pre-trained encoders to achieve the decoupling of portrait motion information and identity in videos. To do so, implicit representation is adopted to encode motion information and is employed as control signals in the animation phase. By leveraging the power of stable video diffusion as the main building block, we carefully design adapter layers to inject control signals into the denoising unet through attention mechanisms. These bring spatial richness of details and temporal consistency. HunyuanPortrait also exhibits strong generalization performance, which can effectively disentangle appearance and motion under different image styles. Our framework outperforms existing methods, demonstrating superior temporal consistency and controllability. Our project is available at https://kkakkkka.github.io/HunyuanPortrait.

肖像动画扩散模型隐式表征可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。