arXiv:2504.04155cs.CL2025-04EMNLP被引 1

构建多语言评测框架,精准诊断大模型在低资源语言中的表现

GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models

  • 设计轻量级多语言评测框架,覆盖7类任务
  • 支持数十至数百种语言,涵盖低资源语种
  • 适合评估跨语言能力,尤其关注非英语场景

大型语言模型在全球快速演进,各地日益将其用于本地语言应用。然而,模型在多样语言环境、特别是低资源语言中的评估已成为学术界和产业界的重大挑战。现有评测体系过度聚焦英语及少数高资源语言,难以反映模型在真实多语言场景下的表现。为此,我们提出GlotEval,一个轻量级的海量多语言评测框架。该框架支持机器翻译、文本分类、摘要生成、开放式生成、阅读理解、序列标注和内在评估共7项关键任务,覆盖数十到数百种语言。GlotEval强调统一多语言基准、语言特异性提示模板以及非英语中心的机器翻译,可精准诊断模型在不同语言环境中的优劣势。一项多语言翻译案例研究验证了其在多语言及语言特定评估中的适用性。

原文摘要 · Abstract (English)

Large language models (LLMs) are advancing at an unprecedented pace globally, with regions increasingly adopting these models for applications in their primary language. Evaluation of these models in diverse linguistic environments, especially in low-resource languages, has become a major challenge for academia and industry. Existing evaluation frameworks are disproportionately focused on English and a handful of high-resource languages, thereby overlooking the realistic performance of LLMs in multilingual and lower-resource scenarios. To address this gap, we introduce GlotEval, a lightweight framework designed for massively multilingual evaluation. Supporting seven key tasks (machine translation, text classification, summarization, open-ended generation, reading comprehension, sequence labeling, and intrinsic evaluation), spanning over dozens to hundreds of languages, GlotEval highlights consistent multilingual benchmarking, language-specific prompt templates, and non-English-centric machine translation. This enables a precise diagnosis of model strengths and weaknesses in diverse linguistic contexts. A multilingual translation case study demonstrates GlotEval's applicability for multilingual and language-specific evaluations.

多语言评估大模型评测低资源语言语言多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。