通过解耦音频与表情,实现情感可控的逼真说话头视频生成
EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion
- 分离音频中的唇动与表情信息,提升生成控制力
- 长时生成中保持表情稳定,性能超越现有方法
- 适合需要高情感表达的虚拟人、短视频生成场景
扩散模型已推动说话头生成技术发展,但仍面临表现力不足、控制性差及长时间生成不稳定的问题。本文提出 EmotiveTalk 框架,首先设计视觉引导的音频信息解耦(V-AID)方法,生成与唇动和表情对齐的音频表征;具体地,在 V-AID 中引入基于扩散的共言语时间扩展(Di-CTE)模块,在多源情绪条件约束下生成表情相关表征。随后,提出情感说话头扩散(ETHD)主干网络,包含表情解耦注入(EDI)模块,可自动从参考图像中解耦表情并融合目标表情信息,显著提升生成表现力。实验表明,EmotiveTalk 能生成富有表现力的说话头视频,在长时生成中保持情绪可控性与稳定性,性能优于现有方法。
原文摘要 · Abstract (English)
Diffusion models have revolutionized the field of talking head generation, yet still face challenges in expressiveness, controllability, and stability in long-time generation. In this research, we propose an EmotiveTalk framework to address these issues. Firstly, to realize better control over the generation of lip movement and facial expression, a Vision-guided Audio Information Decoupling (V-AID) approach is designed to generate audio-based decoupled representations aligned with lip movements and expression. Specifically, to achieve alignment between audio and facial expression representation spaces, we present a Diffusion-based Co-speech Temporal Expansion (Di-CTE) module within V-AID to generate expression-related representations under multi-source emotion condition constraints. Then we propose a well-designed Emotional Talking Head Diffusion (ETHD) backbone to efficiently generate highly expressive talking head videos, which contains an Expression Decoupling Injection (EDI) module to automatically decouple the expressions from reference portraits while integrating the target expression information, achieving more expressive generation performance. Experimental results show that EmotiveTalk can generate expressive talking head videos, ensuring the promised controllability of emotions and stability during long-time generation, yielding state-of-the-art performance compared to existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。