用多任务学习从带字幕语音中学习新词发音,提升语音合成准确率。
Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
- 通过多任务学习融合字幕语音数据,自动提取发音知识。
- 对仅在字幕语音中出现的词汇,音素错误率从2.5%降至1.6%。
- 实现更简单,效果媲美复杂流程,适合语音合成研究者。
近期研究证明,从传统流水线前端迁移构建端到端语言前端是可行且有益的。为克服迁移训练数据固定的词汇覆盖局限,已有工作提出利用易获取的带字幕语音作为额外训练源,以获取未覆盖词汇的新发音知识,但该方法依赖辅助语音识别(ASR)模型,实现流程复杂。本文提出一种基于多任务学习(MTL)的替代方案,直接利用带字幕语音作为训练源。实验表明,相较于基线序列到序列(Seq2Seq)前端,所提MTL方法在仅出现在带字幕语音中的词类型上,将音素错误率(PER)从2.5%降低至1.6%,达到与先前方法相当的效果,但实现流程显著简化。
原文摘要 · Abstract (English)
Recent work has shown the feasibility and benefit of bootstrapping an integrated sequence-to-sequence (Seq2Seq) linguistic frontend from a traditional pipeline-based frontend for text-to-speech (TTS). To overcome the fixed lexical coverage of bootstrapping training data, previous work has proposed to leverage easily accessible transcribed speech audio as an additional training source for acquiring novel pronunciation knowledge for uncovered words, which relies on an auxiliary ASR model as part of a cumbersome implementation flow. In this work, we propose an alternative method to leverage transcribed speech audio as an additional training source, based on multi-task learning (MTL). Experiments show that, compared to a baseline Seq2Seq frontend, the proposed MTL-based method reduces PER from 2.5% to 1.6% for those word types covered exclusively in transcribed speech audio, achieving a similar performance to the previous method but with a much simpler implementation flow.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。