用多模态数据预训练,提升低资源发音合成效果
Deep Speech Synthesis from Multimodal Articulatory Representations
- 用MRI和肌电图数据做多模态预训练,增强发音到语音的转换能力
- 相比之前方法,语音合成错误率降低36%,在多个指标上表现更优
- 适合研究发音机制或语音合成的学者,尤其关注低资源场景
目前可用于训练深度学习模型的发音数据远少于声学语音数据。为提升低资源条件下发音到声学的合成性能,我们提出一种多模态预训练框架。在基于实时磁共振成像(MRI)和表面肌电图(sEMG)输入的单说话人语音合成任务中,合成语音的可懂度显著提升。例如,与之前的工作相比,使用所提出的迁移学习方法使MRI到语音的合成性能提升36%的词错误率。此外,在三个客观和主观合成质量指标上,多模态预训练模型始终优于单模态基线模型。
原文摘要 · Abstract (English)
The amount of articulatory data available for training deep learning models is much less compared to acoustic speech data. In order to improve articulatory-to-acoustic synthesis performance in these low-resource settings, we propose a multimodal pre-training framework. On single-speaker speech synthesis tasks from real-time magnetic resonance imaging and surface electromyography inputs, the intelligibility of synthesized outputs improves noticeably. For example, compared to prior work, utilizing our proposed transfer learning methods improves the MRI-to-speech performance by 36% word error rate. In addition to these intelligibility results, our multimodal pre-trained models consistently outperform unimodal baselines on three objective and subjective synthesis quality metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。