arXiv:2502.12829cs.CL2025-02ACL被引 19

首个面向哈萨克语的多任务评估集,填补中亚语言研究空白

KazMMLU: Evaluating Language Models on Kazakh, Russian, and Regional Knowledge of Kazakhstan

  • 构建2.3万道题的双语评测集,覆盖从STEM到人文的多学科知识
  • 顶尖模型在哈萨克语上表现不佳,准确率远低于高资源语言
  • 适合关注低资源语言、多语种模型与中亚文化技术落地的研究者

尽管哈萨克斯坦人口达两千万,其语言与文化在自然语言处理领域仍严重缺位。尽管大型语言模型(LLMs)全球持续进步,但针对哈萨克语的进展有限,表现为专用模型和基准评估稀缺。为此,我们提出KazMMLU——首个专为哈萨克语设计的MMLU风格数据集。该数据集包含23,000道题目,涵盖多个教育层次,包括科学、技术、工程、数学(STEM)、人文学科与社会科学,均来自真实教学材料,并经母语者及教育工作者人工验证。其中10,969题为哈萨克语,12,031题为俄语,体现哈萨克斯坦双语教育体系与本地知识背景。我们对Llama-3.1、Qwen-2.5、GPT-4与DeepSeek V3等主流多语言模型进行评估,结果显示即使表现最佳的模型在哈萨克语与俄语上的性能仍显著落后,远未达到高资源语言水平。这一发现凸显了低资源语言间的巨大性能差距。我们希望该数据集能推动面向哈萨克语的LLM研究与发展。数据与代码将在论文录用后公开。

原文摘要 · Abstract (English)

Despite having a population of twenty million, Kazakhstan's culture and language remain underrepresented in the field of natural language processing. Although large language models (LLMs) continue to advance worldwide, progress in Kazakh language has been limited, as seen in the scarcity of dedicated models and benchmark evaluations. To address this gap, we introduce KazMMLU, the first MMLU-style dataset specifically designed for Kazakh language. KazMMLU comprises 23,000 questions that cover various educational levels, including STEM, humanities, and social sciences, sourced from authentic educational materials and manually validated by native speakers and educators. The dataset includes 10,969 Kazakh questions and 12,031 Russian questions, reflecting Kazakhstan's bilingual education system and rich local context. Our evaluation of several state-of-the-art multilingual models (Llama-3.1, Qwen-2.5, GPT-4, and DeepSeek V3) demonstrates substantial room for improvement, as even the best-performing models struggle to achieve competitive performance in Kazakh and Russian. These findings underscore significant performance gaps compared to high-resource languages. We hope that our dataset will enable further research and development of Kazakh-centric LLMs. Data and code will be made available upon acceptance.

多语言模型低资源语言哈萨克语评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。