arXiv:2410.13341cs.LGstat.ML2024-10ICLR被引 50

大模型当裁判评估新模型,效果受限于裁判本身水平。

Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data

  • 用少量高质量标签修正大模型裁判的偏见
  • 若裁判不如被评模型,标注量至少需翻倍
  • 适合追求低成本评估但不依赖模型裁判的研究者

高质量标注在快速发展的机器学习生态中日益成为瓶颈。为避免高昂的人工标注成本,可扩展的评估方法成为研究重点。许多研究尝试用强模型替代人工标注作为评估工具,但此类方法引入自偏好等偏差,扭曲模型比较结果。当前新兴的去偏工具通过少量高质量标签修正大量模型判断,但在理论上存在根本限制:当裁判模型性能不如被评估模型时,任何去偏方法都无法将所需真实标注数量减少超过一半。本文揭示了在评估前沿(即新模型可能优于裁判)场景下,大模型作为裁判的严重局限性。实证分析进一步表明,实际中可节省的样本量远低于理论上限。研究还提供了关于去偏方法的新见解,并指明未来改进方向。

原文摘要 · Abstract (English)

High quality annotations are increasingly a bottleneck in the explosively growing machine learning ecosystem. Scalable evaluation methods that avoid costly annotation have therefore become an important research ambition. Many hope to use strong existing models in lieu of costly labels to provide cheap model evaluations. Unfortunately, this method of using models as judges introduces biases, such as self-preferencing, that can distort model comparisons. An emerging family of debiasing tools promises to fix these issues by using a few high quality labels to debias a large number of model judgments. In this paper, we study how far such debiasing methods, in principle, can go. Our main result shows that when the judge is no more accurate than the evaluated model, no debiasing method can decrease the required amount of ground truth labels by more than half. Our result speaks to the severe limitations of the LLM-as-a-judge paradigm at the evaluation frontier where the goal is to assess newly released models that are possibly better than the judge. Through an empirical evaluation, we demonstrate that the sample size savings achievable in practice are even more modest than what our theoretical limit suggests. Along the way, our work provides new observations about debiasing methods for model evaluation, and points out promising avenues for future work.

模型评估大模型裁判去偏方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。