arXiv:2605.04208cs.CL2026-05

评测19个大模型在43种加纳语言上的零样本翻译能力,发现性能普遍不稳。

Nsanku: Evaluating Zero-Shot Translation Performance of LLMs for Ghanaian Languages

  • 构建涵盖43种加纳语言的零样本翻译基准,每语言300句对
  • Gemini-2.5-flash表现最佳(平均26.88分),但无模型在所有语言中兼具高精度与高一致性
  • 开源评估框架可扩展,助力非洲低资源语言NLP研究

大型语言模型(LLMs)在高资源语言上表现出色,但在低资源非洲语言上的表现尚不明确。本文提出Nsanku,一个系统性基准,评估19个开源与专有LLM在43种加纳语言与英语之间的零样本机器翻译性能。数据来自YouVersion圣经平台,每语言300句对。采用双自动指标:BLEU和字符n元语法F分数(chrF),并加入平均准确率与跨语言一致性维度。结果表明,gemini-2.5-flash综合得分最高(26.88,其中BLEU 24.60,chrF 29.16),其次为claude-sonnet-4-5(24.87)、gpt-4.1(23.20)。开源模型中,kimi-k2-instruct-0905领先(20.87)。一致性分析显示,无一模型与语言同时进入高绩效-高一致象限,表明当前模型尚不可靠用于加纳语大规模翻译。Siwu语言得分最高(25.73),Nkonya最低(11.65)。Nsanku是迄今最全面的加纳语言LLM翻译评估,且为公开可扩展的社区基础设施。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated impressive multilingual capabilities for well-resourced languages, yet their performance on low-resource African languages remains poorly understood and largely unevaluated. This paper presents Nsanku, a systematic benchmark that evaluates the zero-shot machine translation performance of 19 open-weight and proprietary LLMs across 43 Ghanaian languages paired with English. Evaluation sentences were sourced from the YouVersion Bible platform, providing 300 sentence pairs per language. Two complementary automatic metrics are employed: Bilingual Evaluation Understudy (BLEU) and Character n-gram F-Score (chrF), alongside an average accuracy score and a cross-language consistency dimension. Nsanku represents the most comprehensive LLM translation evaluation for Ghanaian languages conducted to date. Results show that gemini-2.5-flash achieves the highest overall average score of 26.88 (BLEU: 24.60, chrF: 29.16), followed by claude-sonnet-4-5 at 24.87 (BLEU: 22.46, chrF: 27.28) and gpt-4.1 at 23.20 (BLEU: 21.15, chrF: 25.24). Among open-weight models, kimi-k2-instruct-0905 leads at an average score of 20.87. A critical finding from the consistency analysis is that no model and no language reached the Leaders quadrant of high performance and high consistency simultaneously, indicating that current LLMs are not yet reliably usable for Ghanaian language translation at scale. Siwu achieved the highest per-language average score at 25.73 while Nkonya scored lowest at 11.65. Nsanku establishes a publicly available, community-extensible evaluation infrastructure for African language NLP research.

大模型评测非洲语言零样本翻译多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。