首个专为意大利语设计的LLM评测基准,支持生成与多选任务。
Evalita-LLM: Benchmarking Large Language Models on Italian
- 全用原生意大利语任务,避免翻译偏差和文化偏见。
- 包含生成式与选择题任务,支持更自然的模型交互。
- 多提示评估机制,提升评测公平性,适合本地化研究者使用。
我们介绍了 Evalita-LLM,一个专为评估大型语言模型(LLMs)在意大利语任务上的表现而设计的新基准。其创新特点包括:(i) 所有任务均为原生意大利语,避免了从意大利语翻译带来的问题及潜在文化偏见;(ii) 除传统多项选择题外,还包含生成式任务,使与LLM的交互更自然;(iii) 所有任务均通过多个提示进行评估,降低模型对特定提示的敏感性,实现更公平、客观的评价。我们提出一种迭代方法,通过一组用于开发的LLMs验证候选任务与提示。报告了基准开发阶段的实验结果,并提供了多个先进LLM的性能统计数据。
原文摘要 · Abstract (English)
We describe Evalita-LLM, a new benchmark designed to evaluate Large Language Models (LLMs) on Italian tasks. The distinguishing and innovative features of Evalita-LLM are the following: (i) all tasks are native Italian, avoiding issues of translating from Italian and potential cultural biases; (ii) in addition to well established multiple-choice tasks, the benchmark includes generative tasks, enabling more natural interaction with LLMs; (iii) all tasks are evaluated against multiple prompts, this way mitigating the model sensitivity to specific prompts and allowing a fairer and objective evaluation. We propose an iterative methodology, where candidate tasks and candidate prompts are validated against a set of LLMs used for development. We report experimental results from the benchmark's development phase, and provide performance statistics for several state-of-the-art LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。