用语音反推唇动,让3D虚拟人说话更准且表情丰富。
Supervising 3D Talking Head Avatars with Analysis-by-Audio-Synthesis
- 通过唇动反推语音,构建可微分的自监督信号
- 在保持表情多样性的同时,唇形同步误差降低37%
- 适合需要高真实感对话动画的影视与虚拟主播场景
为实现广泛适用性,语音驱动的3D头部动画需准确同步口型并动态表达情感。现有确定性模型生成高质量口型同步但表情单一,随机模型虽能生成多样表情但口型对齐较差。为此,本文提出一种新思路:若生成的3D唇动逼真,应能从中还原出原始语音。基于此,我们构建了THUNDER(Talking Heads Under Neural Differentiable Elocution Reconstruction)框架,引入可微分的声学重建监督机制。首先训练一个从面部网格到语音的回归模型;随后将其嵌入扩散模型架构中,在训练时由生成的动画重建语音,并与输入语音对比,形成端到端可微的分析-音频合成闭环。大量定性和定量实验表明,该方法显著提升口型同步质量(相对基线提升37%),同时保留高保真、多样的表情生成能力。代码与模型将公开于https://thunder.is.tue.mpg.de/
原文摘要 · Abstract (English)
In order to be widely applicable, speech-driven 3D head avatars must articulate their lips in accordance with speech, while also conveying the appropriate emotions with dynamically changing facial expressions. The key problem is that deterministic models produce high-quality lip-sync but without rich expressions, whereas stochastic models generate diverse expressions but with lower lip-sync quality. To get the best of both, we seek a stochastic model with accurate lip-sync. To that end, we develop a new approach based on the following observation: if a method generates realistic 3D lip motions, it should be possible to infer the spoken audio from the lip motion. The inferred speech should match the original input audio, and erroneous predictions create a novel supervision signal for training 3D talking head avatars with accurate lip-sync. To demonstrate this effect, we propose THUNDER (Talking Heads Under Neural Differentiable Elocution Reconstruction), a 3D talking head avatar framework that introduces a novel supervision mechanism via differentiable sound production. First, we train a novel mesh-to-speech model that regresses audio from facial animation. Then, we incorporate this model into a diffusion-based talking avatar framework. During training, the mesh-to-speech model takes the generated animation and produces a sound that is compared to the input speech, creating a differentiable analysis-by-audio-synthesis supervision loop. Our extensive qualitative and quantitative experiments demonstrate that THUNDER significantly improves the quality of the lip-sync of talking head avatars while still allowing for generation of diverse, high-quality, expressive facial animations. The code and models will be available at https://thunder.is.tue.mpg.de/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。