arXiv:2505.23290cs.SDcs.CV2025-05CVPR被引 8

解决语音动画中相似音节混淆问题,提升面部动作自然度

Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation

  • 引入可插拔语义解耦模块,分离语音特征中的语义信息
  • 实验表明能显著减少相似音节的平均化现象,提升唇形精度
  • 适用于多种语音驱动面部动画模型,增强生成自然性

在3D语音驱动面部动画生成中,现有方法普遍使用预训练自监督语音模型作为编码器。但由于语言中存在大量发音相似但口型不同的音节(近似同音词),这些音节在自监督语音特征空间中容易产生显著耦合,导致后续唇部运动生成出现平均化效应。为此,本文提出一种即插即用的语义解耦模块Wav2Sem。该模块提取整个音频序列对应的语义特征,利用新增的语义信息对特征空间中的音频编码进行解耦,从而获得更具表现力的音频特征。在多个语音驱动模型上的大量实验表明,Wav2Sem模块能有效解耦音频特征,显著缓解发音相似音节在口型生成中的平均化问题,提升面部动画的精确性与自然性。代码已开源:https://github.com/wslh852/Wav2Sem.git。

原文摘要 · Abstract (English)

In 3D speech-driven facial animation generation, existing methods commonly employ pre-trained self-supervised audio models as encoders. However, due to the prevalence of phonetically similar syllables with distinct lip shapes in language, these near-homophone syllables tend to exhibit significant coupling in self-supervised audio feature spaces, leading to the averaging effect in subsequent lip motion generation. To address this issue, this paper proposes a plug-and-play semantic decorrelation module-Wav2Sem. This module extracts semantic features corresponding to the entire audio sequence, leveraging the added semantic information to decorrelate audio encodings within the feature space, thereby achieving more expressive audio features. Extensive experiments across multiple Speech-driven models indicate that the Wav2Sem module effectively decouples audio features, significantly alleviating the averaging effect of phonetically similar syllables in lip shape generation, thereby enhancing the precision and naturalness of facial animations. Our source code is available at https://github.com/wslh852/Wav2Sem.git.

语音动画语义解耦面部建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。