首个面向波斯语大模型的多领域评测基准,揭示现有模型性能不足一半。
FarsEval-PKBETS: A new diverse benchmark for evaluating Persian large language models
- 构建覆盖医学、法律等10个领域的4000道波斯语文本题
- 三款主流模型平均准确率低于50%,多数题答错
- 融合波斯语语言文化特性,适合评估本地化大模型
针对资源丰富的语言如英语,大语言模型(LLM)的评估研究已相当充分,但波斯语等语言的研究仍显不足。本文提出FarsEval-PKBETS基准,是FarsEval项目中用于评估波斯语大语言模型的子集。该基准包含4000个问题与答案,涵盖多项选择、简答和描述性回答等多种形式,覆盖医学、法律、宗教、波斯语知识、百科、人类偏好、社会常识、伦理与偏见、文本生成及尊重他人权利等多个领域。基准设计融入波斯语的语言学、文化和伊朗本地特征。为确保题目对当前模型具有挑战性,使用Llama3-70B、PersianMind和Dorna三个模型进行测试,其平均准确率低于50%,表明这些模型在完全正确回答的问题上少于一半。结果表明,当前语言模型距离有效理解波斯语复杂语境仍有显著差距。
原文摘要 · Abstract (English)
Research on evaluating and analyzing large language models (LLMs) has been extensive for resource-rich languages such as English, yet their performance in languages such as Persian has received considerably less attention. This paper introduces FarsEval-PKBETS benchmark, a subset of FarsEval project for evaluating large language models in Persian. This benchmark consists of 4000 questions and answers in various formats, including multiple choice, short answer and descriptive responses. It covers a wide range of domains and tasks,including medicine, law, religion, Persian language, encyclopedic knowledge, human preferences, social knowledge, ethics and bias, text generation, and respecting others' rights. This bechmark incorporates linguistics, cultural, and local considerations relevant to the Persian language and Iran. To ensure the questions are challenging for current LLMs, three models -- Llama3-70B, PersianMind, and Dorna -- were evaluated using this benchmark. Their average accuracy was below 50%, meaning they provided fully correct answers to fewer than half of the questions. These results indicate that current language models are still far from being able to solve this benchmark
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。