arXiv:2602.05150cs.CL2026-02ACL被引 5

首个基于希腊语母语内容的多任务评测基准,填补了希腊语大模型评估空白。

GreekMMLU: A Native-Sourced Multitask Benchmark for Evaluating Language Models in Greek

  • 从希腊学术、职业和政府考试中采集21,805道原生多选题
  • 80多个模型评测显示顶尖模型与开源模型间性能差距显著
  • 适合研究希腊语NLP或需本地化语言模型评估的团队使用

大语言模型(LLMs)通常在包含希腊语的多语言语料上训练,但针对希腊语的可靠评估基准仍十分有限。现有数据集多为英文机器翻译而来,难以体现希腊语的语言与文化特征。本文提出GreekMMLU,一个面向希腊语大规模多任务语言理解的原生来源评测基准,包含45个学科领域的21,805道多选题,按新定义的主题分类体系组织,并标注了从基础到专业考试级别的教育难度。所有题目均源自或由希腊语母语者编写,来自学术、职业及政府考试。我们公开发布16,857个样本,保留4,948个样本用于私有排行榜,以实现稳健且防污染的评估。对超过80个开源与闭源模型的评测揭示了前沿模型与开源模型之间、希腊语适配模型与通用多语言模型之间的显著性能差距。最后,系统分析了模型规模、适配策略与提示工程等因素对性能的影响,为提升希腊语大模型能力提供了洞见。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are commonly trained on multilingual corpora that include Greek, yet reliable evaluation benchmarks for Greek-particularly those based on authentic, native-sourced content-remain limited. Existing datasets are often machine-translated from English, failing to capture Greek linguistic and cultural characteristics. We introduce GreekMMLU, a native-sourced benchmark for massive multitask language understanding in Greek, comprising 21,805 multiple-choice questions across 45 subject areas, organized under a newly defined subject taxonomy and annotated with educational difficulty levels spanning primary to professional examinations. All questions are sourced or authored in Greek from academic, professional, and governmental exams. We publicly release 16,857 samples and reserve 4,948 samples for a private leaderboard to enable robust and contamination-resistant evaluation. Evaluations of over 80 open- and closed-source LLMs reveal substantial performance gaps between frontier and open-weight models, as well as between Greek-adapted models and general multilingual ones. Finally, we provide a systematic analysis of factors influencing performance-including model scale, adaptation, and prompting-and derive insights for improving LLM capabilities in Greek.

希腊语多任务评测语言模型本土化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。