arXiv:2409.07255cs.CV2024-09被引 7

让说话头像随情绪自然变化,支持精细调控与单次生成。

EMOdiffhead: Continuously Emotional Control in Talking Head Generation via Diffusion

  • 用音频+表情向量引导扩散模型生成情感视频
  • 支持情绪类别和强度的精细控制,实现单次生成
  • 克服情感数据少、多样性差的问题,适合影视动画应用

语音驱动肖像动画任务旨在利用身份图像和语音音频生成说话头像视频。现有方法多关注口型同步与视频质量,却较少研究情感驱动的生成。情感可控性对生成富有表现力的真实动画至关重要。为此,我们提出EMOdiffhead,一种新型情感说话头像生成方法,不仅能精细控制情绪类别与强度,还支持单次生成。基于FLAME 3D模型在表情建模上的线性特性,我们采用DECA方法提取表情向量,并结合音频引导扩散模型生成具有精确口型同步与丰富情感表达的视频。该方法不仅从非情感数据中学习丰富的面部信息,还能有效生成情感视频,克服了情感数据多样性不足、背景信息匮乏及非情感数据缺乏情感细节等局限。大量实验与用户研究证明,本方法在性能上优于现有情感肖像动画方法。

原文摘要 · Abstract (English)

The task of audio-driven portrait animation involves generating a talking head video using an identity image and an audio track of speech. While many existing approaches focus on lip synchronization and video quality, few tackle the challenge of generating emotion-driven talking head videos. The ability to control and edit emotions is essential for producing expressive and realistic animations. In response to this challenge, we propose EMOdiffhead, a novel method for emotional talking head video generation that not only enables fine-grained control of emotion categories and intensities but also enables one-shot generation. Given the FLAME 3D model's linearity in expression modeling, we utilize the DECA method to extract expression vectors, that are combined with audio to guide a diffusion model in generating videos with precise lip synchronization and rich emotional expressiveness. This approach not only enables the learning of rich facial information from emotion-irrelevant data but also facilitates the generation of emotional videos. It effectively overcomes the limitations of emotional data, such as the lack of diversity in facial and background information, and addresses the absence of emotional details in emotion-irrelevant data. Extensive experiments and user studies demonstrate that our approach achieves state-of-the-art performance compared to other emotion portrait animation methods.

说话头像情感控制扩散模型音频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。