评测大模型生成能力,推动无参考自动评估发展
Overview of the NTCIR-18 Automatic Evaluation of LLMs (AEOLLM) Task
- 聚焦生成任务,鼓励无参考评估方法
- 涵盖对话、摘要等4类子任务,覆盖多样生成场景
- 4支团队提交48次运行结果,验证方法有效性
本文综述了NTCIR-18自动评估大语言模型(AEOLLM)任务。随着大语言模型在学术界和工业界广泛应用,如何有效评估其能力成为关键且仍具挑战性的问题。现有评估方法分为人工评估(成本高)与自动评估(受限于题型多为选择题、评价标准依赖参考文本)。为推动自动评估创新,我们提出面向生成任务的AEOLLM任务,倡导无参考评估方法,并设置对话生成、文本扩展、摘要生成及非事实性问答等多样化子任务,全面测试评估方法。今年共收到来自4个团队的48次运行结果。本文将分别介绍任务背景、数据集、评估指标与结果。
原文摘要 · Abstract (English)
In this paper, we provide an overview of the NTCIR-18 Automatic Evaluation of LLMs (AEOLLM) task. As large language models (LLMs) grow popular in both academia and industry, how to effectively evaluate the capacity of LLMs becomes an increasingly critical but still challenging issue. Existing methods can be divided into two types: manual evaluation, which is expensive, and automatic evaluation, which faces many limitations including task format (the majority belong to multiple-choice questions) and evaluation criteria (occupied by reference-based metrics). To advance the innovation of automatic evaluation, we propose the AEOLLM task which focuses on generative tasks and encourages reference-free methods. Besides, we set up diverse subtasks such as dialogue generation, text expansion, summary generation and non-factoid question answering to comprehensively test different methods. This year, we received 48 runs from 4 teams in total. This paper will describe the background of the task, the data set, the evaluation measures and the evaluation results, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。