arXiv:2603.29244cs.CLcs.LG2026-03

构建首个覆盖十种非洲低资源语言的多模态数据集,推动本土语音技术发展。

The Thiomi Dataset: A Large-Scale Multimodal Corpus for Low-Resource African Languages

  • 通过社区众包平台收集超60万条文本与38.5万段音频
  • 在斯瓦希里语上实现3.24%词错误率,提升61%相对性能
  • 适合非洲语言研究、语音合成及低资源语种技术开发者

我们提出Thiomi数据集,一个涵盖十种非洲语言的大规模多模态语料库,涉及四个语言家族:东非的斯瓦希里语、基库尤语、坎巴语、基梅鲁语、卢奥语、马赛语、基普西吉斯语、索马里语;西非的沃洛夫语;以及西中非的富拉尼语。数据集包含超过60.1万条经审核的句子级文本标注和超过38.5万段音频录音,由超过100名贡献者通过专用社区数据采集平台完成。为验证其有效性,我们训练并评估了自动语音识别(ASR)、机器翻译(MT)和文本转语音(TTS)模型,建立了所有语言的基准。最佳ASR系统在斯瓦希里语(Common Voice)上达到3.24%词错误率(WER),将先前学术最优(SOTA)从8.3%降至3.24%(绝对降低5.1个百分点,相对降低61%),索马里语达到4.3% WER。数据集将发布于HuggingFace。本文描述了数据采集平台、质量保障流程及基准实验,并讨论其对非洲语言技术基础设施的意义。

原文摘要 · Abstract (English)

We present the Thiomi Dataset, a large-scale multimodal corpus spanning ten African languages across four language families: Swahili, Kikuyu, Kamba, Kimeru, Luo, Maasai, Kipsigis, Somali (East Africa); Wolof (West Africa); and Fulani (West/Central Africa). The dataset contains over 601,000 approved sentence-level text annotations and over 385,000 audio recordings, collected through a dedicated community data collection platform involving over 100 contributors. To validate the dataset's utility, we train and evaluate ASR, MT, and TTS models, establishing baselines across all languages. Our best ASR system achieves 3.24% WER on Swahili (Common Voice), reducing prior academic SOTA from 8.3% to 3.24% (5.1 percentage point absolute, 61% relative reduction), and 4.3% WER on Somali. The dataset will be published on HuggingFace. We describe the collection platform, quality assurance workflows, and baseline experiments, and discuss implications for African language technology infrastructure.

多模态非洲语言语音识别数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。