首个针对非洲语种埃菲克语的端到端语音合成研究,填补低资源语言技术空白。
Towards Digital Preservation of Efik: TTS for a Low-Resource African Language

- 构建2632条语音的单人语料库,对比四种神经模型在低资源下的表现。
- MMS-TTS获3.80分最高评分,但音调错误仍存在,长句生成较稳定。
- 为非洲声调语言语音合成提供可复现基准,适合低资源语言研究者参考。
埃菲克语是尼日利亚东南部约300万第二语言使用者和150万母语者的声调语言,在语音合成研究中仍严重缺位。本文首次开展埃菲克语端到端文本转语音研究,构建包含2632个发音片段、总计三小时的单说话人语料库,并在低资源条件下对比评估四种神经模型(VITS、MMS-TTS、SpeechT5、Orpheus-TTS)。母语者通过MOS、Nat-MOS和A-MOS进行评价。MMS-TTS取得最高MOS得分3.80 ± 0.63,长句生成更稳定,但音调错误依然存在;其他模型则表现出更大的音调与韵律不一致问题。研究结果为非洲声调语言语音合成提供了可复现基线,凸显了更大语料库与声调感知建模的必要性。
原文摘要 · Abstract (English)
Efik, a tonal language spoken by about 3 million second language speakers and 1.5 million native speakers in Southeastern Nigeria, remains underrepresented in speech synthesis research. We present the first documented end-to-end text-to-speech study for Efik, introducing a curated single speaker corpus of 2,632 utterances totaling three hours and a comparative evaluation of four neural models (VITS, MMS-TTS, SpeechT5, and Orpheus-TTS) under low resource conditions. Native speakers evaluated the systems using MOS, Nat-MOS, and A-MOS. MMS-TTS achieved the highest MOS of 3.80 +/- 0.63 and produced more stable long form speech, though tonal errors persisted. Other models showed greater tonal and prosodic inconsistencies. These results provide a reproducible baseline and highlight the need for larger corpora and tone aware modeling for tonal African languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。