分离姿态与表情,实现高保真可控的肖像动画生成
DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion Representations
- 用显式变换表示姿态,隐式编码表示表情
- 在扩散模型中分别注入姿态与表情信号,提升控制精度
- 支持仅改表情或仅改姿态的精准编辑,适合影视特效应用
单图驱动的肖像动画是长期挑战。现有基于扩散模型的方法虽能生成逼真动画,但难以实现头姿与面部表情的高保真解耦控制,限制了仅修改表情或仅修改姿态的应用。为此,我们提出DeX-Portrait,通过显式全局变换表示姿态、隐式潜在码表示表情,实现解耦驱动的肖像动画生成。首先设计运动训练器,学习精确的姿势与表情编码器;其次采用双分支条件机制将姿态变换注入扩散模型,通过交叉注意力注入表情潜码;最后设计渐进式混合无分类器引导,增强身份一致性。实验表明,该方法在动画质量与解耦可控性上均优于当前最优基线。
原文摘要 · Abstract (English)
Portrait animation from a single source image and a driving video is a long-standing problem. Recent approaches tend to adopt diffusion-based image/video generation models for realistic and expressive animation. However, none of these diffusion models realizes high-fidelity disentangled control between the head pose and facial expression, hindering applications like expression-only or pose-only editing and animation. To address this, we propose DeX-Portrait, a novel approach capable of generating expressive portrait animation driven by disentangled pose and expression signals. Specifically, we represent the pose as an explicit global transformation and the expression as an implicit latent code. First, we design a powerful motion trainer to learn both pose and expression encoders for extracting precise and decomposed driving signals. Then we propose to inject the pose transformation into the diffusion model through a dual-branch conditioning mechanism, and the expression latent through cross attention. Finally, we design a progressive hybrid classifier-free guidance for more faithful identity consistency. Experiments show that our method outperforms state-of-the-art baselines on both animation quality and disentangled controllability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。