arXiv:2411.07127cs.CLcs.LG2024-11ICLR被引 26

无需标准答案,用新指标评估大模型生成判断的质量。

Benchmarking LLMs' Judgments with No Gold Standard

  • 提出GEM指标,通过生成模型估算响应间互信息来评估生成质量。
  • 在人工标注数据上与GPT-4o Examiner表现相当,且抗改写/扩写干扰更强。
  • 构建GRE-bench评测集,专评大模型学术审稿能力,避免数据泄露。

我们提出GEM(生成式互信息估计器),一种用于评估大语言模型(LLMs)语言生成能力的新指标,尤其适用于生成有信息量的判断,且无需依赖黄金标准参考。GEM通过生成模型估算候选响应与参考响应之间的互信息,而不要求参考为黄金标准。在人工标注数据集上的实验表明,GEM与人类评分的相关性达到当前最优水平,优于所有基线方法,且对重述或扩写等策略性操纵更具鲁棒性。此外,我们构建了GRE-bench(生成评审评测基准),基于GEM评估大模型生成高质量学术论文审稿意见的能力。由于采用每年持续更新的开源论文与审稿数据,GRE-bench有效规避了数据污染问题。我们展示了多个主流LLM在ICLR2023数据集上的审稿能力表现。

原文摘要 · Abstract (English)

We introduce the GEM (Generative Estimator for Mutual Information), an evaluation metric for assessing language generation by Large Language Models (LLMs), particularly in generating informative judgments, without the need for a gold standard reference. GEM broadens the scenarios where we can benchmark LLM generation performance-from traditional ones, like machine translation and summarization, where gold standard references are readily available, to subjective tasks without clear gold standards, such as academic peer review. GEM uses a generative model to estimate mutual information between candidate and reference responses, without requiring the reference to be a gold standard. In experiments on a human-annotated dataset, GEM demonstrates competitive correlations with human scores compared to the state-of-the-art GPT-4o Examiner, and outperforms all other baselines. Additionally, GEM is more robust against strategic manipulations, such as rephrasing or elongation, which can artificially inflate scores under a GPT-4o Examiner. We also present GRE-bench (Generating Review Evaluation Benchmark) which evaluates LLMs based on how well they can generate high-quality peer reviews for academic research papers. Because GRE-bench is based upon GEM, it inherits its robustness properties. Additionally, GRE-bench circumvents data contamination problems (or data leakage) by using the continuous influx of new open-access research papers and peer reviews each year. We show GRE-bench results of various popular LLMs on their peer review capabilities using the ICLR2023 dataset.

大模型评估无标准评测学术审稿

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。