用扩散模型分离情绪,生成更真实有情感的3D人脸动画。
EmoDiffusion: Enhancing Emotional 3D Facial Animation with Latent Diffusion Models
- 在隐空间中用双变分自编码器分离上半脸与嘴部动作
- 在3D-BEF数据集上实现比基线高12%的感知质量
- 适合需要精细情绪表达的影视动画与虚拟人开发
语音驱动的3D人脸动画旨在生成与语音内容及情感细节同步的逼真人脸表情,广泛应用于多媒体领域。然而,以往方法常忽视情感表达或难以有效分离情感与语音内容。为此,我们提出EmoDiffusion,通过分离语音中的不同情绪来生成丰富的3D情感表情。该方法采用两个变分自编码器(VAEs)分别生成上半脸区域和嘴部区域,学习更精细的面部序列表征。不同于传统方法将扩散模型直接关联音频与表情序列,我们在隐空间中进行扩散过程。此外,引入情感适配器以更准确评估上半脸运动。由于动画行业缺乏高质量3D情感说话脸数据,我们借助动画专家指导,使用iPhone LiveLinkFace采集面部表情,构建了创新的3D blendshape情感说话脸数据集(3D-BEF),用于训练网络。大量实验与感知评估验证了方法的有效性,结果表明其在生成真实且情感丰富的面部动画方面具有显著优势。
原文摘要 · Abstract (English)
Speech-driven 3D facial animation seeks to produce lifelike facial expressions that are synchronized with the speech content and its emotional nuances, finding applications in various multimedia fields. However, previous methods often overlook emotional facial expressions or fail to disentangle them effectively from the speech content. To address these challenges, we present EmoDiffusion, a novel approach that disentangles different emotions in speech to generate rich 3D emotional facial expressions. Specifically, our method employs two Variational Autoencoders (VAEs) to separately generate the upper face region and mouth region, thereby learning a more refined representation of the facial sequence. Unlike traditional methods that use diffusion models to connect facial expression sequences with audio inputs, we perform the diffusion process in the latent space. Furthermore, we introduce an Emotion Adapter to evaluate upper face movements accurately. Given the paucity of 3D emotional talking face data in the animation industry, we capture facial expressions under the guidance of animation experts using LiveLinkFace on an iPhone. This effort results in the creation of an innovative 3D blendshape emotional talking face dataset (3D-BEF) used to train our network. Extensive experiments and perceptual evaluations validate the effectiveness of our approach, confirming its superiority in generating realistic and emotionally rich facial animations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。