构建首个覆盖2700+语言的多语言词汇能力评测基准
ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models
- 设计8个不同难度任务,评估模型词汇理解与生成能力
- 覆盖2700+语言,远超现有基准的语言广度
- 揭示主流模型在低资源语言上的严重能力短板
现有大语言模型评测基准主要覆盖高/中资源语言,且集中于高层次推理与生成任务。然而大量证据表明,绝大多数世界3800多种书面语言中,大语言模型缺乏基本语言能力。我们提出ChiKhaPo,包含8个不同难度的子任务,用于评估生成模型的词汇理解与生成能力。该基准利用现有词典、单语数据及双语语料,覆盖2700+语言,两个子任务的语言覆盖率超越现有任何基准。我们进一步验证6个顶尖模型在该基准上表现不佳,并分析影响性能的因素,包括语系、资源丰富度、任务类型及理解/生成方向。通过ChiKhaPo,我们希望推动大语言模型的海量多语言评测。
原文摘要 · Abstract (English)
Existing benchmarks for large language models (LLMs) are largely restricted to high- or mid-resource languages, and often evaluate performance on higher-order tasks in reasoning and generation. However, plenty of evidence points to the fact that LLMs lack basic linguistic competence in the vast majority of the world's 3800+ written languages. We introduce ChiKhaPo, consisting of 8 subtasks of varying difficulty designed to evaluate the lexical comprehension and generation abilities of generative models. ChiKhaPo draws on existing lexicons, monolingual data, and bitext, and provides coverage for 2700+ languages for 2 subtasks, surpassing any existing benchmark in terms of language coverage. We further show that 6 SOTA models struggle on our benchmark, and discuss the factors contributing to performance scores, including language family, language resourcedness, task, and comprehension versus generation directions. With ChiKhaPo, we hope to enable and encourage the massively multilingual benchmarking of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。