arXiv:2409.12126cs.CL2024-09NeurIPS被引 24

测试大模型在无语言知识前提下的语言推理能力,发现闭源模型远超开源模型。

Linguini: A benchmark for language-agnostic linguistic reasoning

  • 构建覆盖75种极低资源语言的语义推理测试集,题干含解题所需全部信息
  • 所有模型准确率均低于25%,顶尖闭源模型达24.05%,开源模型仅8.84%
  • 适合研究多语言推理、评估模型泛化能力的研究者参考

我们提出一个新基准,用于衡量语言模型在不依赖预存语言知识的前提下进行语言推理的能力。该测试包含从国际语言奥林匹克竞赛语料库中提取的894道题目,分属160个问题,覆盖75种(主要是)极低资源语言。为在本基准上取得高准确率,模型无需事先掌握目标语言知识,因为所有解题所需信息均在上下文中给出。我们发现,尽管所有分析模型的准确率均低于25%,但开放模型与闭源模型之间存在显著差距:表现最好的专有模型准确率为24.05%,而表现最好的开源模型仅为8.84%。

原文摘要 · Abstract (English)

We propose a new benchmark to measure a language model's linguistic reasoning skills without relying on pre-existing language-specific knowledge. The test covers 894 questions grouped in 160 problems across 75 (mostly) extremely low-resource languages, extracted from the International Linguistic Olympiad corpus. To attain high accuracy on this benchmark, models don't need previous knowledge of the tested language, as all the information needed to solve the linguistic puzzle is presented in the context. We find that, while all analyzed models rank below 25% accuracy, there is a significant gap between open and closed models, with the best-performing proprietary model at 24.05% and the best-performing open model at 8.84%.

语言推理多语言低资源评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。