arXiv:2412.12417cs.CLcs.AI2024-12中稿 · AAAI

为8种非洲低资源语言构建百万级语料库,提升大模型表现

Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments

  • 构建8种非洲语言百万级新语料,覆盖1.6亿母语者
  • 微调使模型平均提升5.6%,文化适配可增3.0%
  • 适合关注公平性、多语言AI的研究者与开发者

大型语言模型(LLMs)在多种任务中表现优异,但非英语语言尤其是非洲本土语言仍存在显著差距。本文为8种低资源非洲语言(阿姆哈拉语、班巴拉语、伊博语、塞佩迪语、绍纳语、塞索托语、茨瓦纳语、科萨语)创建约100万字的人工翻译基准数据,覆盖超过1.6亿使用者。这些基准涵盖Winogrande及MMLU中的大学医学、临床知识和病毒学三个部分。基于此,报告了现有SOTA LLM在英文与非洲语言间的性能差距。通过分析400多个微调模型,探索了高质量数据微调(使用LLM作为标注者)、跨语言迁移和文化适配调整等方法。关键发现包括:微调平均提升5.6%(高质量数据比低质量数据高5.4%),跨语言迁移带来2.9%的平均增益,文化适配问题直接提升3.0%。相关数据、翻译和代码已公开,支持更具包容性的语言技术发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown remarkable performance across various tasks, yet significant disparities remain for non-English languages, and especially native African languages. This paper addresses these disparities by creating approximately 1 million human-translated words of new benchmark data in 8 low-resource African languages, covering a population of over 160 million speakers of: Amharic, Bambara, Igbo, Sepedi (Northern Sotho), Shona, Sesotho (Southern Sotho), Setswana, and Tsonga. Our benchmarks are translations of Winogrande and three sections of MMLU: college medicine, clinical knowledge, and virology. Using the translated benchmarks, we report previously unknown performance gaps between state-of-the-art (SOTA) LLMs in English and African languages. Finally, using results from over 400 fine-tuned models, we explore several methods to reduce the LLM performance gap, including high-quality dataset fine-tuning (using an LLM-as-an-Annotator), cross-lingual transfer, and cultural appropriateness adjustments. Key findings include average mono-lingual improvements of 5.6% with fine-tuning (with 5.4% average mono-lingual improvements when using high-quality data over low-quality data), 2.9% average gains from cross-lingual transfer, and a 3.0% out-of-the-box performance boost on culturally appropriate questions. The publicly available benchmarks, translations, and code from this study support further research and development aimed at creating more inclusive and effective language technologies.

多语言非洲语言模型微调公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。