arXiv:2501.06715cs.CLcs.AI2025-01被引 4

首个针对乌克兰语大模型推理能力的综合性评测基准。

ZNO-Eval: Benchmarking reasoning capabilities of large language models in Ukrainian

  • 基于乌克兰真实高考题构建多类型题目数据集。
  • GPT-4o在语言和常识推理中表现最优,数学题中Gemini Pro与GPT-4 Turbo领先。
  • 揭示了现有模型在乌克兰语与数学任务上的能力短板。

随着大语言模型在文本理解与生成之外的应用增多,评估其能力与局限变得至关重要。尽管近年取得进展,但多数研究聚焦英语,其他语言仍被忽视,导致乌克兰语模型的推理与鲁棒性评估尤为困难。本文提出ZNO-Eval基准,基于乌克兰标准化教育测评系统——外部独立评估(ZNO)与全国多学科测试的真实考题,涵盖单选、多选、匹配及开放问答,覆盖乌克兰语、数学、历史、地理等多领域。对GPT-3.5-Turbo、GPT-4o、GPT-4-Turbo、Mistral Large、Claude 3 Opus、Gemini-1.5 Pro等模型的评测显示,GPT-4o在常识推理与复杂语言任务中表现最佳;而Gemini Pro与GPT-4 Turbo在算术类题目中领先。所有模型在纯文本常识任务(如历史、地理)上接近满分,但在乌克兰语与数学任务上仍有明显差距,凸显为不同语言与场景定制评测基准的重要性。

原文摘要 · Abstract (English)

As the usage of large language models for problems outside of simple text understanding or generation increases, assessing their abilities and limitations becomes crucial. While significant progress has been made in this area over the last few years, most research has focused on benchmarking English, leaving other languages underexplored. This makes evaluating the reasoning and robustness level of language models in Ukrainian particularly challenging. The purpose of this work is to establish a comprehensive benchmark for the reasoning capabilities evaluation of large language models in the Ukrainian language. This paper presents the ZNO-Eval benchmark based on real exam tasks from Ukraine's standardized educational testing system: the External Independent Evaluation and the National Multi-subject Test. With single-answer options, multiple-choice, matching, and open-ended questions from diverse subjects, including Ukrainian language, mathematics, history, and geography, this dataset paves the way toward a thorough analysis of reasoning capabilities across different domains and complexities. Evaluation of several well-known language models, such as GPT-3.5-Turbo, GPT-4o, GPT-4-Turbo, Mistral Large, Claude 3 Opus, and Gemini-1.5 Pro on this benchmark demonstrated the superiority of GPT-4o in both common knowledge reasoning and intricate language tasks. At the same time, Gemini Pro and GPT-4 Turbo excelled in the arithmetic domain, leading in single-answer and open-ended math problems. While all models were close to max performance in text-only common knowledge tasks like history and geography, there still is a gap for Ukrainian language and math, thus highlighting the importance of developing specialized language benchmarks for more accurate assessments of model capabilities and limitations across different languages and contexts.

大模型评测乌克兰语推理能力教育测评

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。