融合几何结构与情感语义,实现高保真情感表情可控生成
Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation

- 用隐式情感特征+显式几何先验联合建模面部动态
- 在保持照片级真实感下实现情绪强度连续控制
- 适合需要精细情绪调节的虚拟人、影视合成场景
音频驱动的情感说话人脸生成旨在合成具有生动面部动态的真实视频。现有方法难以兼顾可控性与视觉保真度。隐式表示虽能捕捉丰富语义,但缺乏结构引导,常导致情绪表达趋均;显式几何方法虽控制性强,却牺牲高频纹理细节。为此,我们提出GemTalk——一种基于扩散模型的框架,融合隐式表示的语义丰富性与显式几何先验的结构精确性。引入视觉引导的音频情绪投影(V-AEP)模块提取隐式情感唇部与表情特征;同时,扩散几何先验生成器(D-GPG)生成与身份相关的混合形状系数作为显式结构先验。关键在于,几何引导的情绪调制(GEM)模块利用这些几何先验重校准隐式特征的幅值,实现对情感表达(尤其是情绪强度)的精确、连续控制,且不损失视觉质量。大量实验表明,GemTalk在照片级真实感与面部情绪动态表现上均达到领先水平。
原文摘要 · Abstract (English)
Audio-driven emotional talking face generation aims to synthesize realistic videos with expressive facial dynamics. However, existing methods struggle to balance controllability and visual fidelity. Although implicit representations capture rich semantics, they lack structural guidance, often resulting in averaged emotional expressions. In contrast, explicit geometric methods offer better control over facial expressions but tend to sacrifice high-frequency texture details. To address it, we propose GemTalk, a diffusion-based framework that combines the semantic richness of implicit representations with the structural precision of explicit geometric priors. We introduce a Vision-guided Audio Emotion Projection (V-AEP) module to extract implicit emotional lip and expression features. At the same time, a Diffusion-based Geometric Priors Generator (D-GPG) generates identity-aware blendshape coefficients as explicit structural priors. Crucially, our Geometry-guided Emotion Modulation (GEM) module leverages these geometric priors to recalibrate the magnitude of implicit features, enabling precise, continuous control over emotional expressions, especially emotion intensity, without sacrificing visual quality. Extensive experiments show GemTalk achieves superior performance in photo-realism, and facial emotional dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。