arXiv:2601.01847cs.CV2026-01被引 4

用3D高斯点实现情绪化且带风格的音频驱动人脸动画

ESGaussianFace: Emotional and Stylized Audio-Driven Facial Animation via 3D Gaussian Splatting

  • 基于3D高斯泼溅重建场景,高效生成三维一致视频
  • 引入情感引导注意力机制,提升不同情绪下的面部细节还原度
  • 支持情绪与风格双重控制,适合影视/虚拟主播等应用

当前大多数音频驱动人脸动画研究聚焦于中性表情生成。尽管部分工作探索了情绪音频驱动的视频生成,但如何高效生成兼具情绪表达与风格特征的高质量说话头视频仍是重大挑战。本文提出ESGaussianFace,一种创新的情绪化与风格化音频驱动人脸动画框架。该方法利用3D Gaussian Splatting重建3D场景并渲染视频,确保生成结果具有高效率与三维一致性。我们设计了情感-音频引导的空间注意力机制,有效融合情感特征与音频内容特征,使模型能更准确地重建不同情绪状态下的面部细节。为实现基于情感与风格特征的3D高斯点形变,引入两个3D高斯形变预测器。此外,提出多阶段训练策略,分步学习角色的唇部动作、情绪变化与风格特征。大量实验表明,本方法在唇部运动准确性、表情变化丰富性及风格表达能力上均优于现有最先进方法。

原文摘要 · Abstract (English)

Most current audio-driven facial animation research primarily focuses on generating videos with neutral emotions. While some studies have addressed the generation of facial videos driven by emotional audio, efficiently generating high-quality talking head videos that integrate both emotional expressions and style features remains a significant challenge. In this paper, we propose ESGaussianFace, an innovative framework for emotional and stylized audio-driven facial animation. Our approach leverages 3D Gaussian Splatting to reconstruct 3D scenes and render videos, ensuring efficient generation of 3D consistent results. We propose an emotion-audio-guided spatial attention method that effectively integrates emotion features with audio content features. Through emotion-guided attention, the model is able to reconstruct facial details across different emotional states more accurately. To achieve emotional and stylized deformations of the 3D Gaussian points through emotion and style features, we introduce two 3D Gaussian deformation predictors. Futhermore, we propose a multi-stage training strategy, enabling the step-by-step learning of the character's lip movements, emotional variations, and style features. Our generated results exhibit high efficiency, high quality, and 3D consistency. Extensive experimental results demonstrate that our method outperforms existing state-of-the-art techniques in terms of lip movement accuracy, expression variation, and style feature expressiveness.

人脸动画3D高斯情绪生成音频驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。