首个评估阿拉伯语方言能力的基准,覆盖5大方言15K题
DialectalArabicMMLU: Benchmarking Dialectal Capabilities in Arabic and Multilingual Language Models
- 手动翻译3000题至5种阿拉伯方言,构建15K问答对
- 19个模型在方言上表现差异大,通用性仍存明显短板
- 适合关注中东语言模型、多语种公平性研究者使用
我们提出 DialectalArabicMMLU,一个用于评估大语言模型在阿拉伯语方言中表现的新基准。尽管近期阿拉伯语及多语种基准已推动现代标准阿拉伯语(MSA)的评估进展,但日常交流中广泛使用的方言仍被严重忽视。DialectalArabicMMLU 在 MMLU-Redux 框架基础上,通过人工翻译与适配,将3000个多项选择题-答案对转化为五大主要方言(叙利亚、埃及、阿联酋、沙特、摩洛哥),共生成15,000个问答对,涵盖32个学术与职业领域(若包含英语和MSA则达22,000对)。该基准支持任务型与语言学层面的系统评估。我们测试了19个开源阿拉伯语及多语种大模型(参数量1B–13B),报告了各方言间显著性能差异,揭示了方言泛化能力的持续不足。DialectalArabicMMLU 是首个统一、人工校准的阿拉伯语方言理解评估资源,推动更包容的评估体系与未来模型发展。
原文摘要 · Abstract (English)
We present DialectalArabicMMLU, a new benchmark for evaluating the performance of large language models (LLMs) across Arabic dialects. While recently developed Arabic and multilingual benchmarks have advanced LLM evaluation for Modern Standard Arabic (MSA), dialectal varieties remain underrepresented despite their prevalence in everyday communication. DialectalArabicMMLU extends the MMLU-Redux framework through manual translation and adaptation of 3K multiple-choice question-answer pairs into five major dialects (Syrian, Egyptian, Emirati, Saudi, and Moroccan), yielding a total of 15K QA pairs across 32 academic and professional domains (22K QA pairs when also including English and MSA). The benchmark enables systematic assessment of LLM reasoning and comprehension beyond MSA, supporting both task-based and linguistic analysis. We evaluate 19 open-weight Arabic and multilingual LLMs (1B-13B parameters) and report substantial performance variation across dialects, revealing persistent gaps in dialectal generalization. DialectalArabicMMLU provides the first unified, human-curated resource for measuring dialectal understanding in Arabic, thus promoting more inclusive evaluation and future model development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。