arXiv:2507.00999cs.CL2025-07ACL被引 3

首个评估西语各变体的开源排行榜,助力多元语言模型发展

La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America

  • 构建涵盖66个数据集的社区驱动评测平台
  • 覆盖巴斯克语、加泰罗尼亚语等10种西语变体,评测50个模型
  • 倡导少样本测试以降低碳排放,促进研究可复现性

排行榜展现了大型语言模型(LLMs)当前的能力与局限。为推动体现西班牙语社群语言与文化多样性的语言模型发展,我们提出La Leaderboard,这是首个针对西班牙及拉丁美洲各语言和语言变体的开源评估平台。该平台由社区共建,旨在为所有关注西语语言模型开发的研究者建立评估标准。当前版本整合了66个数据集,涵盖巴斯克语、加泰罗尼亚语、加利西亚语以及多种西班牙语变体,展示了50个模型的评估结果。为鼓励其他语言的排行榜建设,我们详细说明了方法论,包括针对不同下游任务选择最优评估设置的指导原则。特别地,我们主张采用比文献中更少的少样本示例,以减少环境影响,并提升研究成果对更广泛研究群体的可访问性与可复现性。

原文摘要 · Abstract (English)

Leaderboards showcase the current capabilities and limitations of Large Language Models (LLMs). To motivate the development of LLMs that represent the linguistic and cultural diversity of the Spanish-speaking community, we present La Leaderboard, the first open-source leaderboard to evaluate generative LLMs in languages and language varieties of Spain and Latin America. La Leaderboard is a community-driven project that aims to establish an evaluation standard for everyone interested in developing LLMs for the Spanish-speaking community. This initial version combines 66 datasets in Basque, Catalan, Galician, and different Spanish varieties, showcasing the evaluation results of 50 models. To encourage community-driven development of leaderboards in other languages, we explain our methodology, including guidance on selecting the most suitable evaluation setup for each downstream task. In particular, we provide a rationale for using fewer few-shot examples than typically found in the literature, aiming to reduce environmental impact and facilitate access to reproducible results for a broader research community.

大模型评测多语言西语社区共建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。