arXiv:2411.19799cs.CL2024-11ICLR被引 84

构建多语言知识评估集,测试大模型在真实语境下的理解能力

INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge

  • 基于本地考试题构建44种语言的问答数据集
  • 包含197,243个问题对,覆盖多种区域知识场景
  • 适合评估模型在真实使用环境中的多语言表现

大型语言模型在不同语言间的性能差异阻碍了其在多个地区的有效部署,限制了生成式AI在众多社区中的经济与社会价值。然而,除英语外,许多语言缺乏高质量的评估资源,制约了功能性多语言模型的发展。当前多语言评测常通过翻译英文资源实现,忽略了实际应用环境中所依赖的区域与文化知识。本文构建了一个涵盖44种书面语言的综合性评估套件INCLUDE,包含197,243个来自本地考试资料的问答对,聚焦区域知识与推理能力,旨在评估多语言大模型在真实部署语境中的表现。

原文摘要 · Abstract (English)

The performance differential of large language models (LLM) between languages hinders their effective deployment in many regions, inhibiting the potential economic and societal value of generative AI tools in many communities. However, the development of functional LLMs in many languages (\ie, multilingual LLMs) is bottlenecked by the lack of high-quality evaluation resources in languages other than English. Moreover, current practices in multilingual benchmark construction often translate English resources, ignoring the regional and cultural knowledge of the environments in which multilingual systems would be used. In this work, we construct an evaluation suite of 197,243 QA pairs from local exam sources to measure the capabilities of multilingual LLMs in a variety of regional contexts. Our novel resource, INCLUDE, is a comprehensive knowledge- and reasoning-centric benchmark across 44 written languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed.

多语言知识评估大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。