首个针对拉脱维亚语和吉里马语的LLM评测,揭示模型在低资源语言中的表现差距。
LAG-MMLU: Benchmarking Frontier LLM Understanding in Latvian and Giriama
- 用母语者标注的数据集评测8个顶尖大模型在非英语语言的表现。
- o1模型在英语、拉脱维亚语、吉里马语上0样本得分分别为92.8%、88.8%、70.8%。
- 首次为吉里马语建立基准,凸显低资源语言评估的必要性。
随着大语言模型(LLMs)快速发展,性能评估至关重要。尽管这些模型基于多语言数据训练,其推理能力却主要通过英语数据集评估。因此,亟需使用高质量非英语数据集,尤其是低资源语言(LRLs)构建稳健的评估框架。本研究利用母语者参与构建的大型多任务语言理解(MMLU)子集,对8个前沿大模型在拉脱维亚语和吉里马语上的表现进行评测,其中吉里马语为首次被评测。结果显示,OpenAI的o1模型在所有语言中均表现最优:0样本任务下英语得分为92.8%,拉脱维亚语88.8%,吉里马语70.8%;而Mistral-large(35.6%)和Llama-70B IT(41%)在两种语言中表现均较弱。研究强调了本地化基准与人工评估在推动文化情境化人工智能发展中的重要性。
原文摘要 · Abstract (English)
As large language models (LLMs) rapidly advance, evaluating their performance is critical. LLMs are trained on multilingual data, but their reasoning abilities are mainly evaluated using English datasets. Hence, robust evaluation frameworks are needed using high-quality non-English datasets, especially low-resource languages (LRLs). This study evaluates eight state-of-the-art (SOTA) LLMs on Latvian and Giriama using a Massive Multitask Language Understanding (MMLU) subset curated with native speakers for linguistic and cultural relevance. Giriama is benchmarked for the first time. Our evaluation shows that OpenAI's o1 model outperforms others across all languages, scoring 92.8% in English, 88.8% in Latvian, and 70.8% in Giriama on 0-shot tasks. Mistral-large (35.6%) and Llama-70B IT (41%) have weak performance, on both Latvian and Giriama. Our results underscore the need for localized benchmarks and human evaluations in advancing cultural AI contextualization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。