用情绪感知模型提升少样本3D口型合成的稳定性与真实感
EmoTaG: Emotion-Aware Talking Head Synthesis on Gaussian Splatting with Few-Shot Personalization
- 在FLAME参数空间建模运动,引入几何先验增强稳定性
- 结合音频情感与面部姿态信息,实现更自然的表情变化
- 仅需几秒视频即可个性化,适合影视/虚拟主播场景
基于音频驱动的3D口型合成近年来借助神经辐射场(NeRF)和3D高斯点云(3DGS)取得显著进展。通过利用丰富的预训练先验,少样本方法可仅凭数秒视频完成即时个性化。然而,在表达性面部动作下,现有方法常出现几何不稳与音情不符问题,亟需更有效的情绪感知运动建模。本文提出EmoTaG,一种基于预训练-适配范式的少样本情绪感知3D口型合成框架。核心思想是将运动预测重构为结构化的FLAME参数空间,而非直接变形3D高斯点,从而引入显式几何先验以提升运动稳定性。在此基础上,提出门控残差运动网络(GRMN),从音频中捕捉情感语调,并补充音频中缺失的头部姿态与上半脸线索,实现表达丰富且连贯的运动生成。大量实验表明,EmoTaG在情感表现力、唇形同步、视觉真实感与运动稳定性方面均达到当前最优性能。
原文摘要 · Abstract (English)
Audio-driven 3D talking head synthesis has advanced rapidly with Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS). By leveraging rich pre-trained priors, few-shot methods enable instant personalization from just a few seconds of video. However, under expressive facial motion, existing few-shot approaches often suffer from geometric instability and audio-emotion mismatch, highlighting the need for more effective emotion-aware motion modeling. In this work, we present EmoTaG, a few-shot emotion-aware 3D talking head synthesis framework built on the Pretrain-and-Adapt paradigm. Our key insight is to reformulate motion prediction in a structured FLAME parameter space rather than directly deforming 3D Gaussians, thereby introducing explicit geometric priors that improve motion stability. Building upon this, we propose a Gated Residual Motion Network (GRMN), which captures emotional prosody from audio while supplementing head pose and upper-face cues absent from audio, enabling expressive and coherent motion generation. Extensive experiments demonstrate that EmoTaG achieves state-of-the-art performance in emotional expressiveness, lip synchronization, visual realism, and motion stability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。