arXiv:2606.07167cs.CLcs.AI2026-06

构建首个覆盖26个学科的乌尔都语多任务评测基准,揭示大模型在本土知识上的短板。

UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding

论文配图:UrduMMLU: A Massive Multitask Benchmark for Urdu Language Understanding
图 1 · 摘自论文原文
  • 从本地教辅与考试资料中采集2.6万道乌尔都语多选题,覆盖5大领域
  • 最强模型在乌尔都语上仅达90.34%准确率,开源模型落后超7分
  • 模型在人文类本土内容表现差,尤其比理工科低25至40分

有意义的多语言评估必须在目标语言和教育背景下进行。乌尔都语(使用者超过2.3亿)缺乏基于本土教育资料构建的MMLU式评测基准。我们提出UrduMMLU,一个包含26,431道乌尔都语多选题的基准,覆盖26个学科和5个领域,数据来源为本地乌尔都语题库及公开考试PDF。与翻译生成资源不同,UrduMMLU涵盖标准学术科目和乌尔都语及地区特有内容。通过双人人工标注与严格一致性筛选标记考试来源部分。我们在英语和乌尔都语提示下评估30个大模型,完成60次零样本测试,并对4个开源模型在双语言提示下进行多轮少样本设置评估。Gemini-3.5-Flash表现最佳,分别达到90.20%和90.34%准确率,其余模型均未超过85%。最强开源模型落后7.79和8.92分,多数模型在以乌尔都语为中心的人文类科目上比理工科低25至40分。少样本提示仅带来小幅提升。UrduMMLU表明当前大模型在乌尔都语知识,尤其是区域化内容上仍不均衡。

原文摘要 · Abstract (English)

Meaningful multilingual evaluation must test models in the target language and educational context. Urdu, spoken by more than 230 million people, lacks a broad MMLU-style benchmark built from native educational sources. We introduce UrduMMLU, a benchmark of 26,431 Urdu MCQs across 26 subjects and five domains, collected from native Urdu MCQ banks and public examination PDFs. Unlike translation-based resources, UrduMMLU covers both standard academic subjects and Urdu- and region-specific content. We label the exam-derived portion through dual human annotation with strict consensus filtering. We evaluate 30 LLMs under English and Urdu prompts, yielding 60 zero-shot evaluations, and further evaluate four open-source LLMs under multiple few-shot settings across both prompt languages. Gemini-3.5-Flash performs best, reaching 90.20% and 90.34% accuracy, while no other model exceeds 85%. The strongest open-source model trails by 7.79 and 8.92 points, and many models lose 25 to 40 points on Urdu-centered Humanities subjects compared with STEM. Few-shot prompting yields only modest gains. UrduMMLU shows that Urdu knowledge remains uneven in current LLMs, especially for regionally grounded content.

多语言评测乌尔都语大模型评估本土知识

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。