arXiv:2509.25670cs.SDcs.CV2025-09

针对普通话唇动到语音合成,提升语调准确性和可懂度。

LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning

  • 用跨语言迁移学习复用英语预训练模型,解决音素映射复杂问题。
  • 引入流匹配模型生成声调基频轮廓,提升语调准确性。
  • 适合关注中文语音合成、语调建模的研究者和开发者。

普通话唇动到语音(L2S)合成面临挑战,主要源于复杂的视觉单元到音素映射,以及语调在可懂性中的关键作用。为此,我们提出词汇语调感知的唇动到语音合成(LTA-L2S)。为应对视觉单元到音素的复杂性,模型通过跨语言迁移学习策略适配一个预训练的英语音频-视觉自监督学习(SSL)模型,不仅将英语海量数据中学习到的通用知识迁移到普通话领域,还避免了从头训练该模型的高昂成本。为专门建模词汇语调并提升可懂性,进一步采用流匹配模型生成基频(F0)轮廓,该过程由经过语音识别微调的SSL语音单元引导,包含重要的超音段信息。整体语音质量通过两阶段训练范式进一步提升:第一阶段生成粗略频谱图,第二阶段使用流匹配后处理网络进行精细化优化。大量实验表明,LTA-L2S在语音可懂性和语调准确性方面均显著优于现有方法。

原文摘要 · Abstract (English)

Lip-to-speech (L2S) synthesis for Mandarin is a significant challenge, hindered by complex viseme-to-phoneme mappings and the critical role of lexical tones in intelligibility. To address this issue, we propose Lexical Tone-Aware Lip-to-Speech (LTA-L2S). To tackle viseme-to-phoneme complexity, our model adapts an English pre-trained audio-visual self-supervised learning (SSL) model via a cross-lingual transfer learning strategy. This strategy not only transfers universal knowledge learned from extensive English data to the Mandarin domain but also circumvents the prohibitive cost of training such a model from scratch. To specifically model lexical tones and enhance intelligibility, we further employ a flow-matching model to generate the F0 contour. This generation process is guided by ASR-fine-tuned SSL speech units, which contain crucial suprasegmental information. The overall speech quality is then elevated through a two-stage training paradigm, where a flow-matching postnet refines the coarse spectrogram from the first stage. Extensive experiments demonstrate that LTA-L2S significantly outperforms existing methods in both speech intelligibility and tonal accuracy.

语音合成语调建模跨语言迁移唇动语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。