arXiv:2410.12265cs.CL2024-10AAAI综述被引 1

用自动评审机制高效评估大模型生成质量,成本更低且无偏见。

Auto-PRE: An Automatic and Cost-Efficient Peer-Review Framework for Language Generation Evaluation

  • 模仿人工评审流程,自动筛选具备一致性、相关性与自信心的模型作评价
  • 在摘要、问答、对话任务中表现优于现有方法,成本大幅降低
  • 适合需要大规模自动化评估的LLM研究者和开发者

大语言模型的快速发展凸显了高效可靠评估方法的需求。传统评估常面临成本高、任务形式有限、依赖人工参考及系统性偏差等问题。为此,我们提出Auto-PRE,一种受同行评审启发的自动语言生成评估框架。该框架不依赖人工标注,而是基于三个核心特质——一致性(对应指令理解)、相关性(内容匹配)与自信心(回应可靠性)——自动选取评估模型,覆盖从指令到响应的完整评估流程。在摘要生成、非事实性问答和对话生成三个代表性任务上的实验表明,Auto-PRE达到当前最优性能,同时显著降低评估成本。其结构化与可扩展的设计为实现‘模型作为评判者’的自动化评估提供了关键思路,推动更先进基于LLM的评估体系发展。

原文摘要 · Abstract (English)

The rapid development of large language models (LLMs) has highlighted the need for efficient and reliable methods to evaluate their performance. Traditional evaluation methods often face challenges like high costs, limited task formats, dependence on human references, and systematic biases. To address these limitations, we propose Auto-PRE, an automatic LLM evaluation framework inspired by the peer review process. Unlike previous approaches that rely on human annotations, Auto-PRE automatically selects evaluator LLMs based on three core traits: consistency, pertinence, and self-confidence, which correspond to the instruction, content, and response stages, respectively, and collectively cover the entire evaluation process. Experiments on three representative tasks, including summarization, non-factoid QA, and dialogue generation, demonstrate that Auto-PRE achieves state-of-the-art performance while significantly reducing evaluation costs. Furthermore, the structured and scalable design of our automatic qualification exam framework provides valuable insights into automating the evaluation of LLMs-as-judges, paving the way for more advanced LLM-based evaluation frameworks.

模型评估自动化LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。