arXiv:2507.17578cs.CL2025-07被引 3

用合成语音数据提升非洲低资源语言的语音识别效果

Synthetic Voice Data for Automatic Speech Recognition in African Languages

  • 用大模型生成文本,再通过语音合成构建大规模语料
  • 250小时合成数据+250小时真实数据媲美500小时纯真实数据
  • 为10种非洲语言生成超2500小时语音,成本低于真实数据1%

非洲超过2300种语言的语音技术仍处于缺失状态。本文首次系统评估了大规模合成语音语料在非洲语音识别中的应用。采用三步流程:大语言模型生成文本、语音合成生成声音、语音识别模型微调。八种语言生成的文本可读性评分超过7分中的5分。对豪萨语、多卢语、奇切瓦语三种语言进行评估,共创建超过2500小时合成语音数据,成本不足真实数据的1%。在豪萨语上,使用250小时真实数据与250小时合成数据微调的Wav2Vec-BERT-2.0模型,达到仅使用500小时真实数据的基线性能;而579小时真实数据搭配450至993小时合成数据时表现最优。还进行了性别细分的语音识别性能分析。对于极低资源语言,效果差异显著:奇切瓦语在1:2真实与合成比例下相对词错误率下降约6.5%;多卢语在1:1比例下部分数据集表现相似,但其他数据集未见提升。研究发现需改进评审协议与评估数据准确性。所有数据与模型已公开,鼓励后续研究优化非洲语言的合成数据。

原文摘要 · Abstract (English)

Speech technology remains out of reach for most of the over 2300 languages in Africa. We present the first systematic assessment of large-scale synthetic voice corpora for African ASR. We apply a three-step process: LLM-driven text creation, TTS voice synthesis, and ASR fine-tuning. Eight out of ten languages for which we create synthetic text achieved readability scores above 5 out of 7. We evaluated ASR improvement for three (Hausa, Dholuo, Chichewa) and created more than 2,500 hours of synthetic voice data at below 1% of the cost of real data. Fine-tuned Wav2Vec-BERT-2.0 models trained on 250h real and 250h synthetic Hausa matched a 500h real-data-only baseline, while 579h real and 450h to 993h synthetic data created the best performance. We also present gender-disaggregated ASR performance evaluation. For very low-resource languages, gains varied: Chichewa WER improved about 6.5% relative with a 1:2 real-to-synthetic ratio; a 1:1 ratio for Dholuo showed similar improvements on some evaluation data, but not on others. Investigating intercoder reliability, ASR errors and evaluation datasets revealed the need for more robust reviewer protocols and more accurate evaluation data. All data and models are publicly released to invite further work to improve synthetic data for African languages.

语音识别合成数据非洲语言低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。