arXiv:2606.19792cs.SD2026-06中稿 · INTERSPEECH 2026

预训练模型对新音素生成帮助有限,主要提升语音自然度。

Exploring Pre-training Benefits on Phoneme Addition through Fine-tuning in Speech Synthesis

  • 通过模拟与真实跨语言场景验证音素添加效果
  • 微调需更多数据才能达到与从零训练相当的新音素识别率
  • 预训练核心优势在语音自然度,而非音素扩展能力

迁移学习广泛应用于低资源文本到语音合成。当目标语料包含预训练中未见的音素时,模型需在微调阶段扩充音素库,这一过程称为“音素添加”。然而,预训练模型对已见音素的生成能力是否有助于该过程仍不明确。本研究在两种设置下开展实验:(1) 使用大语言模型生成的音素控制语料进行模拟实验,排除干扰因素;(2) 在真实语音跨语言迁移场景(英语到日语)中验证结果普适性。两种设置下均发现,虽然微调比从零训练更自然,但要达到与从零训练相当的新音素字错误率(PER),所需数据量相当甚至更多。结果表明,预训练主要提升语音自然度,对音素添加贡献有限。

原文摘要 · Abstract (English)

Transfer learning is widely used for low-resource text-to-speech. When the target corpus contains phonemes unseen in pre-training, the model must expand its phoneme inventory during fine-tuning; we call the process "phoneme addition." However, it remains unclear whether the pre-trained ability to generate seen phonemes contributes to this process. This study investigates phoneme addition in two settings: (1) a simulation setup using LLM-generated phoneme-controlled corpora that enables investigation without considering confounding factors, and (2) a real-speech cross-lingual transfer setup (English to Japanese) to validate whether the findings hold in practice. Experiments in both settings showed that while fine-tuning achieved higher naturalness than training from scratch, it required as much or more data to achieve comparable PER for new phonemes. These results indicate that pre-training mainly contributes to naturalness improvement, but offers limited benefit for phoneme addition.

语音合成迁移学习音素添加

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。