面向意大利语的大型语言模型评估新基准,推动社区共建与持续改进。
Challenging the Abilities of Large Language Models in Italian: a Community Initiative
- 汇聚80+学者企业专家,设计20余项任务覆盖语言能力与推理等多维度
- 评测4个开源大模型,揭示其在事实一致性与公平性上的系统性短板
- 强调细粒度指标与统一评估流程,适合关注多语言AI评估的研究者
大型语言模型(LLM)的快速发展已深刻影响自然语言处理领域,但针对非英语语言的系统性评估仍显不足。'挑战意大利语中语言模型能力'(CALAMITA)是由意大利计算语言学协会牵头的大型协作基准项目,不同于传统排行榜,其重点在于方法论:联合来自学术界、产业界和公共部门的80多位贡献者,共同设计、文档化并评估涵盖语言能力、常识推理、事实一致性、公平性、摘要生成、翻译及代码生成等多样任务的集合。通过该过程,我们构建了包含超过20项任务和近100个子任务的基准,并建立支持异构数据集与度量标准的集中式评估流水线。报告了4个开放权重模型的结果,揭示了各能力维度上的系统性优劣势,以及任务特定评估的挑战。除量化结果外,CALAMITA还总结出方法论启示:细粒度且任务代表性的度量标准至关重要,统一的评估流水线不可或缺,社区广泛参与既有优势也有局限。该基准被设计为滚动更新模式,可持续集成新任务与新模型,既是迄今最全面多元的意大利语评估资源,也是可持续、社区驱动评估的框架。我们认为这一组合为其他语言社区提供了包容且严谨的评估实践范本。
原文摘要 · Abstract (English)
The rapid progress of Large Language Models (LLMs) has transformed natural language processing and broadened its impact across research and society. Yet, systematic evaluation of these models, especially for languages beyond English, remains limited. "Challenging the Abilities of LAnguage Models in ITAlian" (CALAMITA) is a large-scale collaborative benchmarking initiative for Italian, coordinated under the Italian Association for Computational Linguistics. Unlike existing efforts that focus on leaderboards, CALAMITA foregrounds methodology: it federates more than 80 contributors from academia, industry, and the public sector to design, document, and evaluate a diverse collection of tasks, covering linguistic competence, commonsense reasoning, factual consistency, fairness, summarization, translation, and code generation. Through this process, we not only assembled a benchmark of over 20 tasks and almost 100 subtasks, but also established a centralized evaluation pipeline that supports heterogeneous datasets and metrics. We report results for four open-weight LLMs, highlighting systematic strengths and weaknesses across abilities, as well as challenges in task-specific evaluation. Beyond quantitative results, CALAMITA exposes methodological lessons: the necessity of fine-grained, task-representative metrics, the importance of harmonized pipelines, and the benefits and limitations of broad community engagement. CALAMITA is conceived as a rolling benchmark, enabling continuous integration of new tasks and models. This makes it both a resource -- the most comprehensive and diverse benchmark for Italian to date -- and a framework for sustainable, community-driven evaluation. We argue that this combination offers a blueprint for other languages and communities seeking inclusive and rigorous LLM evaluation practices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。