让虚拟人脸随情绪连续变化,还能精准对口型。
EmoTalkingGaussian: Continuous Emotion-conditioned Talking Head Synthesis
- 用连续情绪值控制人脸表情,支持动态情感表达。
- 在PSNR、SSIM等指标上优于现有方法,情感与口型更自然。
- 适合需要高情感真实感的虚拟主播、互动应用。
基于3D高斯点云的说话头合成近年来因其高质量图像和实时推理能力受到关注。然而,由于训练仅使用短视频且缺乏面部情绪多样性,生成结果难以表现丰富情感。为此,我们提出一种唇部对齐的情感人脸生成器,并用于训练EmoTalkingGaussian模型。该模型可基于连续情绪值(效价与唤醒度)调控面部表情,同时保持唇动与输入语音同步。为提升野外音频下的唇同步精度,引入自监督学习方法,结合文本转语音网络与视听同步网络。我们在公开视频数据集上验证了该模型,在图像质量(PSNR、SSIM、LPIPS)、情感表达(V-RMSE、A-RMSE、V-SA、A-SA、Emotion Accuracy)和唇同步性(LMD、Sync-E、Sync-C)方面均优于当前最优方法。
原文摘要 · Abstract (English)
3D Gaussian splatting-based talking head synthesis has recently gained attention for its ability to render high-fidelity images with real-time inference speed. However, since it is typically trained on only a short video that lacks the diversity in facial emotions, the resultant talking heads struggle to represent a wide range of emotions. To address this issue, we propose a lip-aligned emotional face generator and leverage it to train our EmoTalkingGaussian model. It is able to manipulate facial emotions conditioned on continuous emotion values (i.e., valence and arousal); while retaining synchronization of lip movements with input audio. Additionally, to achieve the accurate lip synchronization for in-the-wild audio, we introduce a self-supervised learning method that leverages a text-to-speech network and a visual-audio synchronization network. We experiment our EmoTalkingGaussian on publicly available videos and have obtained better results than state-of-the-arts in terms of image quality (measured in PSNR, SSIM, LPIPS), emotion expression (measured in V-RMSE, A-RMSE, V-SA, A-SA, Emotion Accuracy), and lip synchronization (measured in LMD, Sync-E, Sync-C), respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。