arXiv:2603.23654cs.CL2026-03被引 1

为埃塞俄比亚五种主要语言构建高效多语言语音识别模型

Ethio-ASR: Joint Multilingual Speech Recognition and Language Identification for Ethiopian Languages

  • 基于CTC架构联合训练五种埃塞语语音识别模型
  • 平均词错误率30.48%,参数更少却优于现有最强模型
  • 开源模型代码,适合非洲语言研究与低资源语音技术开发者

我们提出Ethio-ASR,一套基于连接时序分类(CTC)的多语言自动语音识别(ASR)模型,联合训练了五种埃塞俄比亚语言:阿姆哈拉语、提格里尼亚语、奥罗莫语、西达马语和沃莱塔语。这些语言分属亚非语系的闪米特、库希特和欧摩蒂克分支,尽管使用者占埃塞人口绝大多数,但在语音技术中仍严重缺失。我们使用新发布的WAXAL语料库,结合多种预训练语音编码器进行训练,并在强基准模型OmniASR上进行评估。最佳模型在WAXAL测试集上达到平均30.48%的词错误率(WER),在参数量显著更少的情况下超越了最优OmniASR模型。我们还深入分析了性别偏差、元音长度与辅音送气对识别错误的影响,以及多语言CTC模型的训练动态。所有模型与代码已公开供研究社区使用。

原文摘要 · Abstract (English)

We present Ethio-ASR, a suite of multilingual CTC-based automatic speech recognition (ASR) models jointly trained on five Ethiopian languages: Amharic, Tigrinya, Oromo, Sidaama, and Wolaytta. These languages belong to the Semitic, Cushitic, and Omotic branches of the Afroasiatic family, and remain severely underrepresented in speech technology despite being spoken by the vast majority of Ethiopia's population. We train our models on the recently released WAXAL corpus using several pre-trained speech encoders and evaluate against strong multilingual baselines, including OmniASR. Our best model achieves an average WER of 30.48% on the WAXAL test set, outperforming the best OmniASR model with substantially fewer parameters. We further provide a comprehensive analysis of gender bias, the contribution of vowel length and consonant gemination to ASR errors, and the training dynamics of multilingual CTC models. Our models and codebase are publicly available to the research community.

多语言识别语音识别低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。