构建200种语言的AI语言能力监测系统,助力低资源语言模型评估
The AI Language Proficiency Monitor -- Tracking the Progress of LLMs on Multilingual Benchmarks
- 整合200种语言的多任务评测,聚焦低资源语言表现
- 覆盖翻译、问答、数学等任务,支持自动更新排行榜
- 适合研究者、开发者及政策制定者追踪模型多语能力
为确保大型语言模型(LLMs)惠及全球语言,必须系统评估其在多语言环境下的能力。我们提出AI语言能力监测系统,全面评估模型在最多200种语言上的表现,尤其关注低资源语言。该基准涵盖翻译、问答、数学与推理等多种任务,使用FLORES+、MMLU、GSM8K、TruthfulQA和ARC等数据集。提供开源、自动更新的排行榜与仪表盘,支持研究人员、开发者与政策制定者识别模型优势与短板。平台不仅排名模型,还提供全球能力分布图与时间趋势分析。本工作补充并扩展了现有多语言基准,旨在推动多语种AI的透明性、包容性与持续进步。系统已上线:https://huggingface.co/spaces/fair-forward/evals-for-every-language。
原文摘要 · Abstract (English)
To ensure equitable access to the benefits of large language models (LLMs), it is essential to evaluate their capabilities across the world's languages. We introduce the AI Language Proficiency Monitor, a comprehensive multilingual benchmark that systematically assesses LLM performance across up to 200 languages, with a particular focus on low-resource languages. Our benchmark aggregates diverse tasks including translation, question answering, math, and reasoning, using datasets such as FLORES+, MMLU, GSM8K, TruthfulQA, and ARC. We provide an open-source, auto-updating leaderboard and dashboard that supports researchers, developers, and policymakers in identifying strengths and gaps in model performance. In addition to ranking models, the platform offers descriptive insights such as a global proficiency map and trends over time. By complementing and extending prior multilingual benchmarks, our work aims to foster transparency, inclusivity, and progress in multilingual AI. The system is available at https://huggingface.co/spaces/fair-forward/evals-for-every-language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。