arXiv:2606.31642cs.CL2026-06

针对6种南部班图语语音识别难题,提出基于语调的渐进式训练方法。

Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition

论文配图:Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition
图 1 · 摘自论文原文
  • 根据语调统计动态调整训练难度,结合门控适配器优化模型
  • 在跨数据集测试中平均词错误率降至28.41%,西茨onga语达23.79%
  • 不同语言需匹配特定模型,建议按语种选型并多数据集验证

南部班图语有超过8000万人使用,但现有基础语音识别模型在零样本条件下词错误率(WER)仍高于100%,难以应用于教育与公共服务。本文针对6种南部班图语,提出一种基于语调条件的渐进式学习框架,融合混合难度评分、由语调统计驱动的门控适配器及分阶段课程训练。在社区语料库上训练,并在NCHLT数据集上测试跨域泛化能力。结果表明,架构与语言间存在明显交互:在宁加语系语言上,W2V-BERT比Whisper低3至4个WER点;而在索托-茨瓦纳语系上Whisper表现更优。采用语调条件的W2V-BERT在所有数据集上平均达到28.41% WER,Xitsonga迁移测试中为23.79%。无单一模型适用于全部6种语言,部署时应按语言选择模型,并在多语料上验证效果。

原文摘要 · Abstract (English)

Southern Bantu languages are spoken by over 80 million people, yet current foundation ASR models still produce zero-shot WER above 100%, which limits practical use in education and public services. We addressed this gap with a tone conditioned curriculum framework for 6 Southern Bantu languages that combined hybrid difficulty scoring, gated adapters driven by tonal statistics and staged curriculum training. We trained on a community corpus and tested transfer to NCHLT to measure robustness beyond matched evaluation. Results revealed clear interactions between architecture and language, with W2V-BERT outperforming Whisper on Nguni languages by 3 to 4 WER points whilst Whisper performed better on Sotho-Tswana languages. W2V-BERT with tone conditioning reached 28.41% average WER across datasets and 23.79% on Xitsonga transfer. No single model suited all 6 languages, so deployment should pair model selection per language with validation across corpora.

语音识别低资源语言语调建模课程学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。