构建首个覆盖伊比利亚语系的通用大模型评估基准,支持持续更新与社区共建。
IberBench: LLM Evaluation on Iberian Languages
- 整合101个数据集,覆盖22类任务,涵盖基础与产业相关自然语言处理
- 23个模型测试显示:工业任务表现普遍低于基础任务,加利西亚语和巴斯克语性能更差
- 支持动态更新、开源代码与公开排行榜,适合多语言研究者和开发者使用
大语言模型(LLMs)在非英语语言上的评估仍面临挑战,尤其因高质量数据稀缺。现有基准多以英语为主,且缺乏语言多样性、忽视产业相关任务、无法动态更新。为此,我们提出IberBench,一个全面可扩展的基准,用于评估伊比利亚半岛及拉丁美洲语言中大模型在基础与产业相关自然语言处理任务上的表现。该基准整合了来自评估活动与最新基准的101个数据集,覆盖22类任务,包括情感分析、毒性检测和摘要生成等。通过支持持续更新与社区提交,并由专家委员会审核,IberBench克服了静态评估的局限性。我们评估了23个参数量从1亿到140亿的LLMs,发现:(i) LLMs在产业任务上表现普遍低于基础任务;(ii) 加利西亚语与巴斯克语下平均性能更低;(iii) 部分任务结果接近随机;(iv) 其他任务虽高于随机但低于共享任务系统。IberBench提供完整开源的评估流程实现,包括数据标准化、增量评估与公开排行榜。
原文摘要 · Abstract (English)
Large Language Models (LLMs) remain difficult to evaluate comprehensively, particularly for languages other than English, where high-quality data is often limited. Existing benchmarks and leaderboards are predominantly English-centric, with only a few addressing other languages. These benchmarks fall short in several key areas: they overlook the diversity of language varieties, prioritize fundamental Natural Language Processing (NLP) capabilities over tasks of industrial relevance, and are static. With these aspects in mind, we present IberBench, a comprehensive and extensible benchmark designed to assess LLM performance on both fundamental and industry-relevant NLP tasks, in languages spoken across the Iberian Peninsula and Ibero-America. IberBench integrates 101 datasets from evaluation campaigns and recent benchmarks, covering 22 task categories such as sentiment and emotion analysis, toxicity detection, and summarization. The benchmark addresses key limitations in current evaluation practices, such as the lack of linguistic diversity and static evaluation setups by enabling continual updates and community-driven model and dataset submissions moderated by a committee of experts. We evaluate 23 LLMs ranging from 100 million to 14 billion parameters and provide empirical insights into their strengths and limitations. Our findings indicate that (i) LLMs perform worse on industry-relevant tasks than in fundamental ones, (ii) performance is on average lower for Galician and Basque, (iii) some tasks show results close to random, and (iv) in other tasks LLMs perform above random but below shared task systems. IberBench offers open-source implementations for the entire evaluation pipeline, including dataset normalization and hosting, incremental evaluation of LLMs, and a publicly accessible leaderboard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。