arXiv:2607.21540cs.CL2026-07

开源非洲语言语音识别模型,用少量标注数据实现高精度跨语言识别。

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

  • 基于w2v-BERT 2.0自监督编码器,采用分步调优策略提升性能。
  • 多语言模型平均词错误率10-13%,接近单语言模型表现。
  • 轻量级语言标识注入机制,单模型可适配多种非洲语言。

我们提出DONDO,一套面向非洲语言的开源、宽松许可的自动语音识别(ASR)基础模型,基于w2v-BERT 2.0自监督语音编码器构建。DONDO包含21个单语言模型和5个多语言模型,覆盖加纳、塞拉利昂、尼日利亚、塞内加尔、肯尼亚和津巴布韦的27种语言变体。模型主要在来自宗教文本的朗读语音上进行微调,这些文本提供广泛、无版权风险且拼写一致的语言数据,弥补了多数非洲语言缺乏标注语音的短板。我们提出两步(部分为三步)学习率退火的微调流程:先以高学习率适应共享多语言模型,再逐步降低以恢复并部分超越强单语言基线。此外,设计了一种轻量级语言条件机制,通过在声学特征前添加一个独热语言标识序列,使单一多语言检查点可在推理时定向至目标语言。在五个多语言家族中,退火后模型平均词错误率(WER)达10-13%,大幅缩小与单语言模型的差距,同时单个检查点支持多种语言。所有模型已在Hugging Face KhayaAI组织下以Apache-2.0许可证(仅需署名)发布,支持自由微调,包括商业用途。保守估计,所覆盖语言的母语使用者约一亿人,若计入第二语言使用者则人数更多。

原文摘要 · Abstract (English)

We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.0 self-supervised speech encoder. DONDO comprises twenty-one monolingual models and five multilingual models spanning twenty-seven language varieties across Ghana, Sierra Leone, Nigeria, Senegal, Kenya and Zimbabwe. Models are fine-tuned primarily on read speech drawn from religious texts, which offer broad, license-clear and orthographically consistent coverage for languages that otherwise lack transcribed audio. We describe a two-step (and, for one family, three-step) learning-rate-annealed fine-tuning procedure that first adapts a shared multilingual model at a high learning rate and then anneals it to recover, and in several cases surpass, strong monolingual baselines. We further describe a lightweight language-conditioning mechanism that injects a one-hot language identity as a sequence of prefix frames prepended to the acoustic features, allowing a single multilingual checkpoint to be steered to a target language at inference. Across the five multilingual families the annealed models reach average word error rates (WER) of 10-13%, closing most of the gap to monolingual models while covering many languages in a single checkpoint. All models are released on the Hugging Face KhayaAI organisation under the Apache-2.0 license (attribution only) so that others may fine-tune them freely, including for commercial use. We provide a conservative estimate that the languages covered are spoken by on the order of one hundred million first-language speakers, and by substantially more when second-language use is included.

语音识别非洲语言多语言开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。