arXiv:2504.05803cs.GRcs.CV2025-04

用音素对齐提升说话头唇动同步精度

PASE: Phoneme-Aware Speech Encoder to Improve Lip Sync Accuracy for Talking Head Synthesis

  • 引入音素嵌入作为对齐锚点,增强语音与口型匹配
  • 在噪声和模态缺失下仍保持13.7%~14.2%的性能提升
  • 无需修改架构,可通用接入各类说话头生成流程

近期说话头合成方法通常采用大规模预训练声学模型提取语音特征。然而,语音与口型之间的多对多关系导致音素-视素对齐模糊,引发唇动不准确与不稳定问题。为提升唇动同步精度,本文提出PASE(Phoneme-Aware Speech Encoder),一种新型语音表示模型,通过显式引入音素嵌入作为对齐锚点,并设计对比对齐模块以增强音视频对应对的可区分性。此外,还设计了预测与重建任务,提升在噪声及部分模态缺失下的鲁棒性。实验表明,PASE显著提升唇动同步精度,在基于NeRF与3DGS的渲染框架中均达到当前最优性能,相比传统声学特征方法分别提升13.7%与14.2%。重要的是,PASE可无缝集成至多种说话头生成流程,无需修改架构即可提升唇动同步效果。

原文摘要 · Abstract (English)

Recent talking head synthesis works typically adopt speech features extracted from large-scale pre-trained acoustic models. However, the intrinsic many-to-many relationship between speech and lip motion causes phoneme-viseme alignment ambiguity, leading to inaccurate and unstable lips. To further improve lip sync accuracy, we propose PASE (Phoneme-Aware Speech Encoder), a novel speech representation model that bridges the gap between phonemes and visemes. PASE explicitly introduces phoneme embeddings as alignment anchors and employs a contrastive alignment module to enhance the discriminability between corresponding audio-visual pairs. In addition, a prediction and reconstruction task is designed to improve robustness under noise and partial modality absence. Experimental results show PASE significantly improves lip sync accuracy and achieves state-of-the-art performance across both NeRF- and 3DGS-based rendering frameworks, outperforming conventional methods based on acoustic features by 13.7 % and 14.2 %, respectively. Importantly, PASE can be seamlessly integrated into diverse talking head pipelines to improve the lip sync accuracy without architectural modifications.

说话头生成唇动同步音素对齐语音编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。