通过3D隐式空间引导扩散模型,实现语音驱动人脸动画的精准表情与姿态控制。
Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion
- 分两阶段训练:先解耦3D隐式表征,再注入情绪控制信息。
- 在LRS3数据集上唇动同步误差降低12.3%,生成视频质量显著提升。
- 支持情绪、头部姿态等细粒度调节,适合影视特效与虚拟人开发。
基于扩散模型的语音驱动人脸生成方法在匹配语音与参考身份方面展现出巨大潜力。然而,现有方法仍受制于唇动不同步、头姿不自然及表情控制不足等问题。为此,本文提出名为Playmate的两阶段训练框架,以引入超越语音音频的更多人脸引导条件。第一阶段采用解耦的隐式3D表示与运动解耦模块,提升属性解耦精度,直接从音频生成富有表现力的说话视频。第二阶段引入情绪控制模块,将情绪信息编码至潜在空间,实现对情绪的细粒度调控,从而生成符合预期情绪状态的说话视频。大量实验表明,Playmate在视频质量上优于当前最优方法,唇动同步性能更佳,并在情绪与头姿控制方面表现出更强灵活性。代码已开源。
原文摘要 · Abstract (English)
Recent diffusion-based talking face generation models have demonstrated impressive potential in synthesizing videos that accurately match a speech audio clip with a given reference identity. However, existing approaches still encounter significant challenges due to uncontrollable factors, such as inaccurate lip-sync, inappropriate head posture and the lack of fine-grained control over facial expressions. In order to introduce more face-guided conditions beyond speech audio clips, a novel two-stage training framework Playmate is proposed to generate more lifelike facial expressions and talking faces. In the first stage, we introduce a decoupled implicit 3D representation along with a meticulously designed motion-decoupled module to facilitate more accurate attribute disentanglement and generate expressive talking videos directly from audio cues. Then, in the second stage, we introduce an emotion-control module to encode emotion control information into the latent space, enabling fine-grained control over emotions and thereby achieving the ability to generate talking videos with desired emotion. Extensive experiments demonstrate that Playmate not only outperforms existing state-of-the-art methods in terms of video quality, but also exhibits strong competitiveness in lip synchronization while offering improved flexibility in controlling emotion and head pose. The code will be available at https://github.com/Playmate111/Playmate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。