测试31种语言下大模型的本地知识能力,发现多数模型表现不佳且语言间差异巨大。
MultiLoKo: a multilingual local knowledge benchmark for LLMs spanning 31 languages
- 构建涵盖31种语言的本地化知识评测集,分主数据集与双向人工翻译集
- 11个主流大模型平均得分低,最佳与最差语言得分差距超20分
- 人工翻译数据比机器翻译更能反映真实语言难度,影响模型排名和性能评估
我们提出MultiLoKo,一个覆盖31种语言的多语言大模型评测基准。该基准包含三个部分:每种语言500个问题的主数据集(本地相关),以及30种非英语到英语和反向的人工翻译数据集;同时提供对应机器生成翻译。数据均分至开发集和盲测分布外测试集。我们对11个标称多语言的基座与对话模型进行测评,分析其平均表现、跨语言性能一致性、语言依赖性及最难语言。结果表明所有模型表现均不理想,平均分低,最佳与最差语言得分相差超过20分。问题语言显著影响模型表现,显示跨语言知识迁移不足。使用本地数据与英文翻译数据可导致最优模型得分差异超20分,显著改变语言难度评估。相比人工翻译,机器翻译虽弱化语言难度排序影响,但大幅降低所有模型评分并改变模型排名。
原文摘要 · Abstract (English)
We present MultiLoKo, a new benchmark for evaluating multilinguality in LLMs covering 31 languages. MultiLoKo consists of three partitions: a main partition consisting of 500 questions per language, separately sourced to be locally relevant to the specific language, and two translated partitions, containing human-authored translations from 30 non-English languages to English and vice versa. For comparison, we also release corresponding machine-authored translations. The data is equally distributed over two splits: a dev split and a blind, out-of-distribution test split. MultiLoKo can be used to study a variety of questions regarding the multilinguality of LLMs as well as meta-questions about multilingual benchmark creation. We compute MultiLoKo scores for 11 base and chat models marketed to be multilingual and study their average performance, their performance parity across languages, how much their ability to answer questions depends on the question language, and which languages are most difficult. None of the models we studied performs well on MultiLoKo, as indicated by low average scores as well as large differences between the best and worst scoring languages. Furthermore, we find a substantial effect of the question language, indicating sub-optimal knowledge transfer between languages. Lastly, we find that using local vs English-translated data can result in differences more than 20 points for the best performing models, drastically change the estimated difficulty of some languages. For using machines instead of human translations, we find a weaker effect on ordering of language difficulty, a larger difference in model rankings, and a substantial drop in estimated performance for all models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。