arXiv:2607.17173cs.CL2026-07

首个针对吉尔吉斯语的大规模语言模型评测基准,填补了低资源语言评估空白。

KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding

  • 构建了原生吉尔吉斯语数据集与人工校对的翻译数据集相结合的评测体系。
  • 26个模型在零样本和少样本下测试,发现英语到吉尔吉斯语迁移效果不一。
  • 公开所有数据与代码,助力未来吉尔吉斯语自然语言研究。

跨语言大语言模型评估仍面临挑战,因多数多语言基准依赖英文数据翻译,常掩盖目标语言的语言文化特异性。这一问题在吉尔吉斯语等低资源语言中尤为突出,可靠原生语料稀缺。本文基于已有吉尔吉斯语评估数据集,首次系统性地开展大规模语言模型在吉尔吉斯语上的评估,推出KyrgyzLLM-Bench评测套件。该套件包含两个原生数据集:KyrgyzMMLU与KyrgyzRC,以及经精心翻译并人工校对的WinoGrande、HellaSwag、BoolQ和TruthfulQA。我们评估了26个开源与闭源模型在零样本与少样本设置下的表现,分析了模型性能、跨语言迁移能力及翻译偏差对评估可靠性的影响。结果显示,在不同模型族与任务中,模型排名在WinoGrande和BoolQ上从英语向吉尔吉斯语有较广泛迁移,而在MMLU上较弱;而HellaSwag则表现出显著的英-吉性能差距,与翻译引发的合理性偏移一致。少样本提示可提升部分开源模型在阅读理解上的表现,但对闭源模型在翻译任务上效果不一。所有数据集、评估代码与模型结果均已公开,并集成至主流多语言评估框架,以支持后续吉尔吉斯语自然语言处理研究。

原文摘要 · Abstract (English)

Evaluating large language models (LLMs) across languages remains challenging, as most multilingual benchmarks rely on translated English datasets, often obscuring linguistic and cultural specificity in the target language. This issue is particularly pronounced for less-resourced languages such as Kyrgyz, where reliable natively authored evaluation data are scarce. Building on previously introduced Kyrgyz-language evaluation datasets, this work reports the first systematic and large-scale evaluation of LLMs in Kyrgyz using the KyrgyzLLM-Bench benchmark suite. KyrgyzLLM-Bench comprises two natively authored datasets$-$KyrgyzMMLU and KyrgyzRC$-$together with carefully translated and manually post-edited versions of WinoGrande, HellaSwag, BoolQ, and TruthfulQA. We evaluate 26 open- and closed-source LLMs under zero-shot and few-shot settings, analyzing model performance, cross-lingual transfer, and the impact of translation artifacts on evaluation reliability. Across families and tasks, model rankings transfer broadly from English to Kyrgyz on WinoGrande and BoolQ, and to a lesser extent on MMLU, while HellaSwag exhibits a substantial English-Kyrgyz performance gap consistent with translation-induced plausibility shifts. Few-shot prompting improves several open-source models on reading comprehension but behaves inconsistently for proprietary models on translated tasks. We publicly release all datasets, evaluation code, and per-model results, and integrate the Kyrgyz tasks into a widely used multilingual evaluation framework to support future research on Kyrgyz NLP.

语言模型低资源语言评测基准吉尔吉斯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。