为19种非洲语言构建高质量开源翻译数据,提升小模型性能。
TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation

- 通过质量筛选与合成数据融合,优化低资源语言训练数据
- 0.8B参数模型超越更大规模的Gemma和Qwen系统
- 适合关注非洲语言AI、小模型高效训练的研究者
人工智能快速发展却未惠及非洲语言,导致数字鸿沟。现有开源大模型在非洲语言翻译上表现不佳,主要因缺乏大规模高质量开源双语数据,制约了小型语言模型(SLMs)的发展。本文提出TranslatePsy-AfriSLM,包含19种撒哈拉以南非洲语言的精选双语数据、专用于非洲语言的合成数据及一系列微调后的SLMs。实证研究表明,统一的质量估计过滤可移除高达96%的训练样本而不降低质量,且过滤后的合成数据在质量-效率权衡中占据最优位置。基于此混合数据微调的TranslatePsy-AfriSLMs仅用0.8B参数,即显著优于更大的TranslateGemma-27B和Qwen3.5-122B-A10B系统。
原文摘要 · Abstract (English)
The rapid progress in Artificial Intelligence has largely bypassed African languages, creating a digital divide that limits AI adoption on the continent. Recent open-source LLMs systematically underperform on African language machine translation, while the lack of large-scale, high-quality, open-source parallel data has constrained the development of competitive small language models (SLMs). We introduce TranslatePsy-AfriSLM, a collection of open-source machine translation resources for 19 Sub-Saharan African languages, including curated parallel data, African-specialized synthetic data, and a family of fine-tuned SLMs. Our empirical study shows that unified quality-estimation filtering removes up to 96% of training tokens without degrading quality, and that filtered synthetic data dominates the quality-efficiency Pareto frontier. Fine-tuned on the resulting data mixture, TranslatePsy-AfriSLMs outperform substantially larger systems, including TranslateGemma-27B and Qwen3.5-122B-A10B, with as few as 0.8B parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。