首个大规模孟加拉语多任务理解评测,覆盖41个领域13万道题
BnMMLU: Measuring Massive Multitask Language Understanding in Bengali
- 构建覆盖41个领域的孟加拉语多任务评测集,含134,375道选择题
- 发现主流模型在推理与应用能力上仍有明显短板,模型规模增益递减
- 提供可复现的评测模板,适合关注低资源语言NLP的研究者
大规模多任务基准推动了语言建模的快速发展,但多数侧重高资源语言如英语,导致孟加拉语严重不足。我们提出BnMMLU,一个全面评估孟加拉语大规模多任务理解能力的基准。BnMMLU涵盖STEM、人文学科、社会科学及通用知识共41个领域,包含134,375个多项选择题-选项对,是迄今最全面的孟加拉语评估套件。数据集通过MathML保留数学内容,并引入BnMMLU-HARD子集,由顶级系统最常出错的问题构成,用于测试难题。我们在11个大语言模型家族的24个模型变体上进行评估,涵盖开源通用/多语言、孟加拉语专用开源及专有模型,覆盖多种参数规模和指令微调设置。采用标准化协议,在两种提示方式(直接式与思维链)和两种上下文设置(零样本与五样本)下报告准确率,结果跨模型家族具有一致性。分析揭示推理与应用能力存在持续差距,且模型规模增益呈亚线性。我们公开数据集与评测模板,支持严谨、可复现的孟加拉语理解评估,推动多语言NLP发展。
原文摘要 · Abstract (English)
Large-scale multitask benchmarks have driven rapid progress in language modeling, yet most emphasize high-resource languages such as English, leaving Bengali underrepresented. We present BnMMLU, a comprehensive benchmark for measuring massive multitask language understanding in Bengali. BnMMLU spans 41 domains across STEM, humanities, social sciences, and general knowledge, and contains 134,375 multiple-choice question-option pairs--the most extensive Bengali evaluation suite to date. The dataset preserves mathematical content via MathML, and includes BnMMLU-HARD, a compact subset constructed from questions most frequently missed by top systems to stress difficult cases. We benchmark 24 model variants across 11 LLM families, spanning open-weights general/multilingual, Bengali-centric open-weights, and proprietary models, covering multiple parameter scales and instruction-tuned settings. We evaluate models under standardized protocols covering two prompting styles (Direct vs. Chain-of-Thought) and two context regimes (0-shot vs. 5-shot), reporting accuracy consistently across families. Our analysis highlights persistent gaps in reasoning and application skills and indicates sublinear returns to scale across model sizes. We release the dataset and evaluation templates to support rigorous, reproducible assessment of Bengali language understanding and to catalyze progress in multilingual NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。