构建非洲语言大模型数据集,推动低资源语言NLP发展
The African Languages Lab: A Collaborative Approach to Advancing Low-Resource African NLP
- 建立40种非洲语言的多模态数据集,含190亿词和1.26万小时语音
- 微调后在31种语言上平均提升23.69点ChrF++和15.34点BLEU
- 培养15名青年研究者,助力本地科研能力建设
尽管非洲语言占全球语言近三分之一,但在现代自然语言处理中仍严重缺位,88%被列为极度低资源或完全忽略。本文提出非洲语言实验室(All Lab),通过系统化数据收集、模型开发与能力建设,填补技术空白。贡献包括:(1) 建立高质量数据采集流程,产出目前最大且经验证的非洲多模态语料库,覆盖40种语言,包含190亿词单语文本和12,628小时对齐语音数据;(2) 实验验证表明,结合微调后,跨31种语言平均提升23.69点ChrF++、0.33点COMET和15.34点BLEU;(3) 构建结构化研究项目,成功培养15名早期科研人员,建立可持续本地能力。与Google Translate对比显示,在部分语言上表现相当,但仍存在需改进领域。
原文摘要 · Abstract (English)
Despite representing nearly one-third of the world's languages, African languages remain critically underserved by modern NLP technologies, with 88\% classified as severely underrepresented or completely ignored in computational linguistics. We present the African Languages Lab (All Lab), a comprehensive research initiative that addresses this technological gap through systematic data collection, model development, and capacity building. Our contributions include: (1) a quality-controlled data collection pipeline, yielding the largest validated African multi-modal speech and text dataset spanning 40 languages with 19 billion tokens of monolingual text and 12,628 hours of aligned speech data; (2) extensive experimental validation demonstrating that our dataset, combined with fine-tuning, achieves substantial improvements over baseline models, averaging +23.69 ChrF++, +0.33 COMET, and +15.34 BLEU points across 31 evaluated languages; and (3) a structured research program that has successfully mentored fifteen early-career researchers, establishing sustainable local capacity. Our comparative evaluation against Google Translate reveals competitive performance in several languages while identifying areas that require continued development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。