为突厥语族构建首个原生多任务语言理解基准,填补高资源语言评估空白。
TUMLU: A Unified and Native Language Understanding Benchmark for Turkic Languages
- 原创设计突厥语族原生多任务评测集,覆盖8种语言11个学科
- 测试多种大模型在不同语言、学科和字母系统下的表现差异
- 开源精简版TUMLU-mini及评估脚本,推动多语言研究
全面评估多任务语言理解能力对提升多语言大模型应用至关重要。然而,高质量母语基准的构建成本高昂,限制了评估数据集的代表性。现有工作多依赖高资源语言机器翻译生成,易引入误差且忽略目标语言的语言文化特性。本文针对语义结构独特、文化特征鲜明的突厥语族,提出首个原生多任务语言理解基准TUMLU。该基准包含阿塞拜疆语、克里米亚鞑靼语、卡拉卡尔帕克语、哈萨克语、鞑靼语、土耳其语、维吾尔语和乌兹别克语共8种语言,涵盖中学至高中水平的11个学科领域。同时推出更简洁、平衡且人工验证过的子集TUMLU-mini。基于此,我们系统评估了Claude、Gemini、GPT和LLaMA等开源与专有大模型的表现,深入分析其在不同语言、学科和文字系统中的性能差异。为促进多语言理解研究,我们公开发布TUMLU-mini及全部评估脚本。
原文摘要 · Abstract (English)
Being able to thoroughly assess massive multi-task language understanding (MMLU) capabilities is essential for advancing the applicability of multilingual language models. However, preparing such benchmarks in high quality native language is often costly and therefore limits the representativeness of evaluation datasets. While recent efforts focused on building more inclusive MMLU benchmarks, these are conventionally built using machine translation from high-resource languages, which may introduce errors and fail to account for the linguistic and cultural intricacies of the target languages. In this paper, we address the lack of native language MMLU benchmark especially in the under-represented Turkic language family with distinct morphosyntactic and cultural characteristics. We propose two benchmarks for Turkic language MMLU: TUMLU is a comprehensive, multilingual, and natively developed language understanding benchmark specifically designed for Turkic languages. It consists of middle- and high-school level questions spanning 11 academic subjects in Azerbaijani, Crimean Tatar, Karakalpak, Kazakh, Tatar, Turkish, Uyghur, and Uzbek. We also present TUMLU-mini, a more concise, balanced, and manually verified subset of the dataset. Using this dataset, we systematically evaluate a diverse range of open and proprietary multilingual large language models (LLMs), including Claude, Gemini, GPT, and LLaMA, offering an in-depth analysis of their performance across different languages, subjects, and alphabets. To promote further research and development in multilingual language understanding, we release TUMLU-mini and all corresponding evaluation scripts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。