arXiv:2606.22274cs.CLcs.AI2026-06

用语音识别扩充西非低资源语言文本,验证了方法可行性

From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongbe and Hausa

论文配图:From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongbe and Hausa
图 1 · 摘自论文原文
  • 用自建语料微调模型,实现高保真音调与变音符号还原
  • 产出6770段转录文本,豪萨语质量达可用水平,丰贝语需后处理
  • 开源数据集、模型与视频目录,支持后续研究

低资源非洲语言缺乏用于语言模型训练的文本语料。本文研究语音识别(ASR)管道能否扩展两种类型迥异的西非语言——丰贝语(声调语言,带变音符号)和豪萨语(非声调语言)的文本资源。针对丰贝语,我们基于12.3小时精选语料微调MMS-300M模型,在ALFFA基准上达到9.48%的词错误率(WER),相比先前44.04%的基线降低78%,并成功保留关键声调变音符号。对于豪萨语,采用已微调的Whisper-Small模型。我们收集1,553个YouTube视频(共236小时),从中选取424个视频(45.49小时)以平衡领域多样性与算力限制,生成6,770段转录内容。人工评估随机抽取的每种语言50段样本,豪萨语平均质量得分57.4/100,丰贝语为36.5/100,表明豪萨语转录已具备构建语料库的可用性,而丰贝语仍需后处理或更优模型。所有数据集、微调模型、转录语料及完整视频目录均按平台规范与伦理准则公开。

原文摘要 · Abstract (English)

Low-resource African languages lack text corpora needed for language model training. We investigate whether ASR pipelines can extend text resources for two typologically distinct West African languages: Fongbe (tonal, diacritic-rich) and Hausa (non-tonal). We fine-tune MMS-300M on a curated 12.3-hour Fongbe dataset, achieving 9.48% WER on the ALFFA benchmark - a 78% relative reduction from the prior 44.04% baseline - while preserving tonal diacritics critical to the language. For Hausa, we apply an existing fine-tuned Whisper-Small model. We catalog 1,553 YouTube videos (236 hours) and process a subset of 424 videos (45.49 hours) selected to balance domain diversity with available computational resources, producing 6,770 transcribed segments. Human evaluation on 50 randomly sampled segments per language shows mean quality scores of 57.4/100 for Hausa and 36.5/100 for Fongbe, indicating that while Hausa transcriptions approach acceptable quality for corpus construction, Fongbe transcriptions require post-processing or improved models for production use. We release the curated dataset, fine-tuned model, transcribed corpus, and full video catalog following platform terms and ethical guidelines.

语音识别低资源语言文本生成数据采集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。