针对波斯语与伊朗文化,构建19个新评测集,评估41个大模型表现。
MELAC: Massive Evaluation of Large Language Models with Alignment of Culture in Persian Language
- 构建19个聚焦伊朗法律、波斯语法等文化的评测数据集
- 在41个主流大模型上测试,揭示跨文化能力差异
- 适合关注非西方文化评估的AI研究者与开发者
随着大语言模型日益融入日常生活,跨语境的质量与可靠性评估变得至关重要。尽管英语领域已有全面基准,但其他语言的评估资源仍严重不足。且多数大模型训练数据源于欧美文化,对非西方文化缺乏理解。为此,本研究聚焦波斯语与伊朗文化,构建19个新评测数据集,涵盖伊朗法律、波斯语法、波斯习语及大学入学考试等内容。基于这些数据集,我们对41个主流大模型进行了评测,旨在弥合该领域在文化和语言评估上的空白。
原文摘要 · Abstract (English)
As large language models (LLMs) become increasingly embedded in our daily lives, evaluating their quality and reliability across diverse contexts has become essential. While comprehensive benchmarks exist for assessing LLM performance in English, there remains a significant gap in evaluation resources for other languages. Moreover, because most LLMs are trained primarily on data rooted in European and American cultures, they often lack familiarity with non-Western cultural contexts. To address this limitation, our study focuses on the Persian language and Iranian culture. We introduce 19 new evaluation datasets specifically designed to assess LLMs on topics such as Iranian law, Persian grammar, Persian idioms, and university entrance exams. Using these datasets, we benchmarked 41 prominent LLMs, aiming to bridge the existing cultural and linguistic evaluation gap in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。