梳理豪萨语和丰贝语的文本语音资源现状与空白
A Survey of Text and Speech Resources for Hausa and Fongbe: Availability, Quality, and Gaps for NLP Development
- 系统整理两类西非语言的公开资源,涵盖文本、语音与模型
- 豪萨语有新闻百科教育多领域文本,丰贝语侧重近期语音数据
- 指出丰贝语缺乏多样文本、豪萨语缺专用语音语料等关键缺口
本综述全面盘点了两种西非语言——豪萨语(约8000万至1亿使用者)和丰贝语(贝宁约200万人使用)——的公开文本与语音资源。这两类语言在资源丰富程度上形成鲜明对比。通过检索学术库、数据平台与网络来源,我们整理了双语语料、单语文本、语音数据集、预训练模型及评估基准。对每项资源,记录其规模、领域覆盖、格式、许可与可访问性。结果显示,豪萨语在新闻、百科、教育等领域拥有较丰富的文本资源;丰贝语虽文本较少,但近年有学术语音采集项目。两类语言均纳入马萨卡内(Masakhane)的命名实体识别与词性标注评测。本文提出任务导向建议,并指出优先补足:丰贝语领域多样文本、豪萨语专用语音语料等关键空白。
原文摘要 · Abstract (English)
This survey provides a comprehensive catalog of publicly available text and speech resources for two West African languages: Hausa, an Afroasiatic language with approximately 80-100 million speakers, and Fongbe, a Niger-Congo language spoken by approximately 2 million people in Benin. These languages represent contrasting cases on the resource availability spectrum. We address the question: \textit{What is the current state of publicly available NLP resources for Hausa and Fongbe, and what gaps remain?} Through systematic search of academic repositories, data platforms, and web sources, we catalog parallel corpora, monolingual text collections, speech datasets, pre-trained models, and evaluation benchmarks. For each resource, we document size, domain coverage, format, licensing, and accessibility. Our findings reveal that Hausa benefits from broader text resource diversity across news, encyclopedic, and educational domains. Fongbe, while having more limited text resources, has been the focus of recent academic speech data collection initiatives. Both languages are represented in Masakhane benchmarks for NER and POS tagging. We provide task-specific recommendations and identify priority gaps including domain-diverse Fongbe text and dedicated Hausa speech corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。