构建南非多语言术语库,助力本地化AI发展
Mafoko: Structuring and Building Open Multilingual Terminologies for South African NLP
- 整合政府与学术机构的零散术语资源,统一格式为可计算数据
- 在英-茨瓦纳语翻译中显著提升大模型准确性与领域一致性
- 开源共享,支持非洲本土AI技术公平发展
南非官方语言缺乏结构化术语数据,严重制约多语言自然语言处理进展。尽管存在大量政府和学术机构的术语列表,但这些资源分散且以非机器可读格式保存,难以用于计算研究。Mafoko通过系统性地聚合、清洗和标准化这些零散资源,构建了开放、可互操作的术语数据集。我们推出了基础版Mafoko数据集,采用非洲中心化的NOODL框架发布。为验证其实际效用,我们将术语集成至检索增强生成(RAG)管道中,实验显示在大型语言模型的英-茨瓦纳语机器翻译任务中,准确率与领域一致性均显著提升。Mafoko为构建稳健、公平的NLP技术提供了可扩展基础,确保南非丰富的语言多样性在数字时代得到体现。
原文摘要 · Abstract (English)
The critical lack of structured terminological data for South Africa's official languages hampers progress in multilingual NLP, despite the existence of numerous government and academic terminology lists. These valuable assets remain fragmented and locked in non-machine-readable formats, rendering them unusable for computational research and development. Mafoko addresses this challenge by systematically aggregating, cleaning, and standardising these scattered resources into open, interoperable datasets. We introduce the foundational Mafoko dataset, released under the equitable, Africa-centered NOODL framework. To demonstrate its immediate utility, we integrate the terminology into a Retrieval-Augmented Generation (RAG) pipeline. Experiments show substantial improvements in the accuracy and domain-specific consistency of English-to-Tshivenda machine translation for large language models. Mafoko provides a scalable foundation for developing robust and equitable NLP technologies, ensuring South Africa's rich linguistic diversity is represented in the digital age.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。