arXiv:2509.03162cs.CL2025-09EMNLP被引 4

首个针对僧伽罗语的多任务语言理解评测集,填补低资源语言评估空白。

SinhalaMMLU: A Comprehensive Benchmark for Evaluating Multitask Language Understanding in Sinhala

  • 专为僧伽罗语设计,覆盖6大领域30个学科的7000+道多选题
  • 最大模型仅达67%准确率,人文类题目表现显著偏低
  • 适合作为低资源语言与文化语境下LLM评估的研究基准

大型语言模型在通用知识和推理能力上表现出色,但其评估长期集中于全球性或英语主导的主题,忽视了低资源语言和文化特异性内容。尽管近年已有多种多语言评测集尝试弥补这一缺口,但多数依赖自动翻译,易引入错误并扭曲原始文化语境。为此,我们提出SinhalaMMLU,首个专为僧伽罗语设计的多项选择问答评测集。该数据集包含超过7,000道题目,覆盖中学至大学教育水平,紧扣斯里兰卡国家课程体系,涵盖六大领域与30个学科,既包括通用学术主题,也包含深厚文化背景的知识。我们在SinhalaMMLU上评估了26个LLM,发现尽管Claude 3.5 sonnet和GPT-4o分别达到67%和62%的最高平均准确率,整体性能仍有限。尤其在人文等文化密集型领域,模型表现明显不足,表明当前LLM在适应低资源与文化特定语境方面仍有巨大提升空间。

原文摘要 · Abstract (English)

Large Language Models (LLMs) demonstrate impressive general knowledge and reasoning abilities, yet their evaluation has predominantly focused on global or anglocentric subjects, often neglecting low-resource languages and culturally specific content. While recent multilingual benchmarks attempt to bridge this gap, many rely on automatic translation, which can introduce errors and misrepresent the original cultural context. To address this, we introduce SinhalaMMLU, the first multiple-choice question answering benchmark designed specifically for Sinhala, a low-resource language. The dataset includes over 7,000 questions spanning secondary to collegiate education levels, aligned with the Sri Lankan national curriculum, and covers six domains and 30 subjects, encompassing both general academic topics and culturally grounded knowledge. We evaluate 26 LLMs on SinhalaMMLU and observe that, while Claude 3.5 sonnet and GPT-4o achieve the highest average accuracies at 67% and 62% respectively, overall model performance remains limited. In particular, models struggle in culturally rich domains such as the Humanities, revealing substantial room for improvement in adapting LLMs to low-resource and culturally specific contexts.

语言模型低资源语言文化理解评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。