通过语音上下文建模,让3D人脸动画更自然流畅。
Learning Phonetic Context-Dependent Viseme for Enhancing Speech-Driven 3D Facial Animation
- 引入语音上下文感知损失,动态调整面部动作权重。
- 在多个指标上优于传统方法,动画更平滑连续。
- 适合做语音驱动人脸动画的研究者和开发者。
语音驱动的3D人脸动画旨在生成与音频同步的逼真面部动作。传统方法主要通过逐帧对齐真实标签来最小化重建误差,但这种逐帧策略难以捕捉面部运动的连续性,导致因音素共现效应而产生抖动和不自然的输出。为此,我们提出一种新颖的语音上下文感知损失,显式建模语音上下文对视觉音素(viseme)转换的影响。通过引入视觉音素共现权重,根据面部动作随时间的动态变化自适应分配重要性,确保动画更平滑且感知一致。大量实验表明,用我们的损失替换传统重建损失后,在定量指标和视觉质量上均有提升。结果凸显了在合成自然语音驱动3D人脸动画中显式建模语音上下文依赖型视觉音素的重要性。
原文摘要 · Abstract (English)
Speech-driven 3D facial animation aims to generate realistic facial movements synchronized with audio. Traditional methods primarily minimize reconstruction loss by aligning each frame with ground-truth. However, this frame-wise approach often fails to capture the continuity of facial motion, leading to jittery and unnatural outputs due to coarticulation. To address this, we propose a novel phonetic context-aware loss, which explicitly models the influence of phonetic context on viseme transitions. By incorporating a viseme coarticulation weight, we assign adaptive importance to facial movements based on their dynamic changes over time, ensuring smoother and perceptually consistent animations. Extensive experiments demonstrate that replacing the conventional reconstruction loss with ours improves both quantitative metrics and visual quality. It highlights the importance of explicitly modeling phonetic context-dependent visemes in synthesizing natural speech-driven 3D facial animation. Project page: https://cau-irislab.github.io/interspeech25/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。