用大模型辅助人类评估,科学设计抽样减少人工成本。
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
- 大模型先全量评分,再对子样本请人复核,分两阶段优化评估效率。
- 提出双重稳健估计器,确保在模型预测不准时仍能准确估算结果。
- 根据大模型预测能力差异动态分配人力,关键类型多留人工审核。
大型语言模型(LLMs)正越来越多地被用作AI系统的自动化评估者,尤其在高风险场景中。尽管专家人工评分成本高且难扩展,而大模型评分快速廉价,但现有方法缺乏严谨设计,仅依赖人与模型评分的一致性指标作为替代依据。本文将大模型评估者的角色从替代转为辅助,提出一种双阶段抽样设计:第一阶段由大模型对所有样本进行评分,第二阶段对部分样本邀请人类评审。我们引入缺失数据领域的双重稳健估计器,利用已知的抽样设计来提升估计稳定性。基于该估计器的渐近方差,我们给出确定人类与大模型样本量的方法,以达到目标检验效能。同时证明,可针对大模型预测能力较弱的评估类型,增加人类评分比例,实现高效研究设计。据我们所知,这是首个系统指导验证基准中保留多少人工监督的研究。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used as automated evaluators of AI systems, including in high-stakes applications. In this role, LLMs are used to generate judgments about the quality, appropriateness, or even safety of model outputs. This approach is motivated by practical constraints. Expert human ratings are costly and difficult to scale, whereas LLM ratings can be produced quickly at low cost. However, current approaches to deploying LLM evaluators are ad hoc, typically limited to reporting agreement metrics between human and LLM judges as a justification for substitution of human ratings, and lack a formal basis for study design. This paper (1) shifts the role of the LLM judge from substitutive to auxiliary, and (2) formulates the LLM-as-a-judge paradigm as one of augmenting human evaluation through a two-stage sampling design, where LLM evaluations are measured for all observations at the first stage and human ratings are partially observed for a subsample at the second stage. We propose to use a doubly robust estimator from the missing data literature, which takes advantage of the robustness property against the prediction model, since the missingness model is known by design. Using the asymptotic variance of this estimator, we propose how sample sizes of human and LLM ratings can be determined to achieve a targeted level of power. We also show that a study can be efficiently designed by allocating more human ratings for types of evaluations where the predictability of LLM ratings is not high. To the best of our knowledge, there is very little guidance on how much human oversight should be retained when validating benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。