arXiv:2410.08928cs.CLcs.AI2024-10被引 35

为21种欧洲语言构建多语言评估基准,验证翻译对模型表现的影响。

Towards Multilingual LLM Evaluation for European Languages

  • 用5个主流基准的翻译版评测40个大模型在21种欧洲语言的表现。
  • 发现翻译服务差异显著影响评估结果,需谨慎选择。
  • 开源全新多语言数据集,助力跨语言模型研究。

大型语言模型(LLM)的兴起彻底改变了自然语言处理领域,但在多种欧洲语言间保持一致且有意义的评估仍具挑战性,尤其受限于缺乏语言对齐的多语言基准。本文提出一种专为欧洲语言设计的多语言评估方法,采用五个广泛使用的基准的翻译版本,评估40个大模型在21种欧洲语言中的表现。贡献包括:分析翻译基准的有效性,评估不同翻译服务的影响,并提供包含新创建数据集的多语言评估框架——EU20-MMLU、EU20-HellaSwag、EU20-ARC、EU20-TruthfulQA 和 EU20-GSM8K。所有基准与评估结果均公开,以推动多语言大模型评估研究。

原文摘要 · Abstract (English)

The rise of Large Language Models (LLMs) has revolutionized natural language processing across numerous languages and tasks. However, evaluating LLM performance in a consistent and meaningful way across multiple European languages remains challenging, especially due to the scarcity of language-parallel multilingual benchmarks. We introduce a multilingual evaluation approach tailored for European languages. We employ translated versions of five widely-used benchmarks to assess the capabilities of 40 LLMs across 21 European languages. Our contributions include examining the effectiveness of translated benchmarks, assessing the impact of different translation services, and offering a multilingual evaluation framework for LLMs that includes newly created datasets: EU20-MMLU, EU20-HellaSwag, EU20-ARC, EU20-TruthfulQA, and EU20-GSM8K. The benchmarks and results are made publicly available to encourage further research in multilingual LLM evaluation.

多语言评估大模型欧洲语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。