融合语音与语言模型提升日语语调标注精度。
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
- 结合语音与音素预训练模型提取多模态特征
- 在日语语调标注上达到94.3%断句准确率
- 适合语音合成与语音标注任务研究者
本文提出一种自动语调标注模型,所生成标签可用于训练可控制语调的文本转语音系统。模型融合自监督学习语音模型(如Whisper编码器)提取的声学特征,以及音素输入的预训练语言模型(如PnG BERT和PL-BERT)获取的语言特征,通过拼接两者特征进行音素级语调标签预测。在日语语调标注实验中,包括重音与语句断点,结果显示联合使用语音与语言基础模型显著优于单一使用。具体表现:重音标签准确率89.8%,高低音重音93.2%,断句索引94.3%。
原文摘要 · Abstract (English)
This paper proposes a model for automatic prosodic label annotation, where the predicted labels can be used for training a prosody-controllable text-to-speech model. The proposed model utilizes not only rich acoustic features extracted by a self-supervised-learning (SSL)-based model or a Whisper encoder, but also linguistic features obtained from phoneme-input pretrained linguistic foundation models such as PnG BERT and PL-BERT. The concatenation of acoustic and linguistic features is used to predict phoneme-level prosodic labels. In the experimental evaluation on Japanese prosodic labels, including pitch accents and phrase break indices, it was observed that the combination of both speech and linguistic foundation models enhanced the prediction accuracy compared to using either a speech or linguistic input alone. Specifically, we achieved 89.8% prediction accuracy in accent labels, 93.2% in high-low pitch accents, and 94.3% in break indices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。