构建100+语言文化常识推理评测集,揭示大模型在低资源语言上的显著短板
Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures
- 由全球350多位研究者手工构建跨141种语言的常识推理数据集
- 高资源语言与低资源语言间准确率差距达68%,凸显模型文化适配不足
- 覆盖五大洲、24种书写系统,含大量本地化习俗与饮食内容
目前几乎不存在覆盖多语言与多文化的大型语言模型评估基准。本文提出Global PIQA,一个由来自65个国家超过350名研究人员手工构建的跨100+语言的参与式常识推理基准。该数据集包含141种语言变体,覆盖五大洲、19个语系和24种书写系统。非平行部分中,超50%的题目涉及本地食物、习俗、传统等文化特异性内容;平行部分将更“文化中立”的常识推理问题翻译成131种语言变体,用于跨语言比较。所有示例均经母语者验证。我们发现,当前顶尖大模型整体表现尚可,但在低资源语言上表现明显下降(平行部分最高达68%准确率差距)。Global PIQA表明,日常常识仍是多数语言文化下大模型亟待提升的关键领域,远超复杂推理与专家知识的讨论范畴。该数据集不仅可用于模型评估,也展现了人类语言所嵌入的文化多样性。
原文摘要 · Abstract (English)
To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we present Global PIQA, a participatory commonsense reasoning benchmark for over 100 languages, constructed by hand by over 350 researchers from over 65 countries around the world. The 141 language varieties in Global PIQA cover five continents, 19 language families, and 24 writing systems. In the non-parallel split of Global PIQA, over 50% of examples reference local foods, customs, traditions, or other culturally-specific elements. In the parallel split, we translate more "culturally agnostic" commonsense reasoning questions into 131 language varieties, for direct cross-lingual comparisons. In both splits, all examples have been verified by native speakers of the languages. We find that state-of-the-art LLMs perform well on Global PIQA in aggregate, but they exhibit weaker performance in lower-resource languages (e.g. up to a 68% accuracy gap between languages in the parallel split). Global PIQA highlights that in many languages and cultures, everyday knowledge remains an area for improvement in LLMs, alongside more widely-discussed capabilities such as complex reasoning and expert knowledge. Beyond its uses for LLM evaluation, Global PIQA provides a glimpse into the wide diversity of cultures in which human language is embedded.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。