arXiv:2603.13793cs.CLcs.AI2026-03被引 2

构建4万+平行语料库,助力加纳低资源语言的AI应用

GhanaNLP Parallel Corpora: Comprehensive Multilingual Resources for Low-Resource Ghanaian Languages

  • 人工采集翻译5种加纳方言与英语的平行句对
  • 共41,513组语料,含结构化元数据确保可用性
  • 支持机器翻译、语音技术及语言保护,适合非洲语言研究者

低资源语言因数字化和结构化语言数据稀缺,面临自然语言处理挑战。为填补这一空白,GhanaNLP倡议开发并整理了41,513组平行句子对,涵盖特维语、芳蒂语、埃韦语、加语和库萨尔语,这些语言在加纳广泛使用但数字空间代表性不足。每组数据均由专业人员完成采集、翻译与标注,并附带标准结构化元数据以保证一致性与可用性。该语料库旨在支持机器翻译、语音技术及语言保护等研究、教育与商业应用。本文详细说明数据集创建方法、结构、预期用途与评估结果,并展示其在Khaya AI翻译引擎中的实际部署。整体工作推动了人工智能普惠化,使非洲语言能获得更包容、可及的语言技术。

原文摘要 · Abstract (English)

Low resource languages present unique challenges for natural language processing due to the limited availability of digitized and well structured linguistic data. To address this gap, the GhanaNLP initiative has developed and curated 41,513 parallel sentence pairs for the Twi, Fante, Ewe, Ga, and Kusaal languages, which are widely spoken across Ghana yet remain underrepresented in digital spaces. Each dataset consists of carefully aligned sentence pairs between a local language and English. The data were collected, translated, and annotated by human professionals and enriched with standard structural metadata to ensure consistency and usability. These corpora are designed to support research, educational, and commercial applications, including machine translation, speech technologies, and language preservation. This paper documents the dataset creation methodology, structure, intended use cases, and evaluation, as well as their deployment in real world applications such as the Khaya AI translation engine. Overall, this work contributes to broader efforts to democratize AI by enabling inclusive and accessible language technologies for African languages.

低资源语言平行语料非洲语言机器翻译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。