用子词融合提升语音合成音素编码器预训练效率
CraBERT: Efficient Phoneme Encoder Pre-Training via Cascade Fusion of Subword Representations for Text-to-Speech

- 通过级联融合子词与音素表示,整合上下文信息
- 仅需1轮预训练即达基线10轮效果,主观评分接近最优
- 适合追求高效音色合成的工业级语音系统研发
本文提出CraBERT,一种用于文本到语音(TTS)的高效音素编码器预训练方法。CraBERT采用级联融合架构与子词-音素对齐算法,将预训练的子词级BERT表示融合至音素级BERT中,引入词汇与句子级先验信息,从而大幅减少音素编码器所需的预训练量。主观听感评估显示,CraBERT在约1个训练周期后达到与现有音素编码器相当的平均意见分(MOS),而对比基线需约10个周期预训练。结果表明,CraBERT可高效学习适用于提升合成语音自然度与韵律表现力的表征。
原文摘要 · Abstract (English)
This paper introduces CraBERT, a pre-trained phoneme encoder (PPEnc) designed for efficient pre-training in text-to-speech (TTS). CraBERT employs a cascade-fusion architecture and a subword-phoneme alignment algorithm to integrate representations from a pre-trained subword-level BERT into a phoneme-level BERT. This design provides prior word- and sentence-level information, reducing the amount of pre-training required by the phoneme encoder. Subjective listening evaluations show that CraBERT achieves MOS values comparable to existing PPEncs after approximately one epoch of pre-training, whereas the baselines in our comparison are pre-trained for approximately ten epochs. These results demonstrate that CraBERT can efficiently learn representations suitable for improving the perceived naturalness and prosody of synthesized speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。