arXiv:2607.00959cs.CV2026-07

用3D高斯点云实现音频驱动的实时情感化人脸合成

GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian Splatting

论文配图:GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian Splatting
图 1 · 摘自论文原文
  • 将情感动画建模为中性到情感的残差变形问题
  • 实现实时渲染,唇动同步准确且情感强度可控
  • 适合需要高保真表情控制的虚拟主播、游戏角色

音频驱动的人脸合成在唇部同步和视觉质量方面已取得显著进展,但在实时约束下生成可调控情感强度的生动虚拟形象仍具挑战。本文提出GaussianEmoTalker,一种基于3D高斯点云的音频驱动实时情感化人脸合成框架。不直接从语音预测最终情感形象,而是将情感动画建模为从中性到情感的残差变形问题。该方法首先构建基于高斯混合形状(GaussianBlendshapes)的身份特异性中性说话空间,提供高保真高斯属性与音素同步的中性运动;随后通过融合网格位移线索、音频特征、情感类别与强度编码,预测情感条件下的残差变形。为融合异构信号,引入空间-音频-情感注意力模块,估算高斯属性的偏移量,实现表达丰富且时间稳定的渲染。大量实验表明,GaussianEmoTalker在视频质量、唇动同步精度、情感可控性及实时渲染性能上均优于近期方法。

原文摘要 · Abstract (English)

Audio-driven talking head synthesis has achieved impressive progress in lip synchronization and visual quality, yet generating expressive emotional avatars with controllable intensity remains challenging, especially under real-time constraints. In this paper, we present GaussianEmoTalker, an audio-driven framework for real-time emotional talking head synthesis based on 3D Gaussian Splatting. Instead of directly predicting the final emotional avatar from speech, we formulate emotional animation as a neutral-to-emotional residual deformation problem. GaussianEmoTalker first constructs an identity-specific neutral talking space with GaussianBlendshapes, which provides high-fidelity Gaussian attributes and phoneme-synchronized neutral motion. It then predicts an emotion-conditioned residual deformation by combining mesh displacement cues, audio features, emotion categories, and intensity encodings. To fuse these heterogeneous signals, we introduce a spatial-audio-emotion attention module that estimates the offsets of Gaussian attributes for expressive and temporally stable rendering. Extensive experiments demonstrate that GaussianEmoTalker achieves competitive video quality, accurate lip synchronization, controllable emotional expression, and real-time rendering compared with recent emotional talking head methods. Our project page is available at https://njust-yang.github.io/GaussianEmoTalker.github.io/

人脸合成3D高斯情感控制实时渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。