arXiv:2511.23370cs.CL2025-11

首个专为非洲语言训练的大规模自监督语音模型,性能显著提升。

Scaling HuBERT for African Languages: From Base to Large and XL

  • 构建了仅用非洲语音数据训练的Large和XL级HuBERT模型。
  • 在撒哈拉以南语言上,大模型显著提升语音识别与语言识别效果。
  • 开源模型权重,适合低资源语言研究与应用开发。

尽管多语言语音处理取得进展,非洲语言在研究与实际系统中仍严重缺失,尤其缺乏强健、可公开获取权重的编码器,在低资源条件下表现良好。自监督学习在此类场景中展现出巨大潜力,但目前公开的针对非洲语音的模型大多仅为基础版(BASE),尚未验证仅基于非洲语音数据训练的大规模编码器是否带来实际效益,以及模型容量与数据构成之间的关系。本工作填补该空白,提出SSA-HuBERT-Large(317M参数)和SSA-HuBERT-XL(964M参数),为首个仅在非洲语音上训练的大型模型,同时包含一个BASE版本。我们通过聚焦撒哈拉以南语言的受控实验,覆盖自动语音识别(ASR)与语言识别(LID)任务,证明更大架构能有效利用大规模音频数据,显著提升性能。

原文摘要 · Abstract (English)

Despite recent progress in multilingual speech processing, African languages remain under-represented in both research and deployed systems, particularly when it comes to strong, open-weight encoders that transfer well under low-resource supervision. Self-supervised learning has proven especially promising in such settings, yet most publicly released models targeting African speech remain at BASE scale, leaving unanswered whether larger encoders, trained exclusively on Africa-centric audio, offer tangible benefits and how model capacity interacts with data composition. This work addresses that gap by introducing SSA-HuBERT-Large (317M parameters) and SSA-HuBERT-XL (964M parameters), the first large models trained solely on African speech, alongside a BASE size counterpart. We release these models as open weights: see https://huggingface.co/collections/Orange/african-speech-foundation-models. By conducting a carefully controlled experimental study focused exclusively on Sub-Saharan languages, covering automatic speech recognition (ASR) and language identification (LID) tasks, we demonstrate that larger architectures significantly improve performance by effectively leveraging large audio datasets.

语音识别自监督学习非洲语言大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。