用历史数据提升生成模型评估的准确性和敏感度
HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data

- 利用历史标注数据校准噪声标签,降低评估偏差
- 通过高精度协变量锚定,显著减少评估方差
- 适用于多轮模型迭代评估,尤其适合资源有限场景
可靠的生成式AI模型依赖专家人工标注来评估输出质量,但这类'黄金'标签成本高且数量有限。因此,组织常采用众包或供应商标注的大量但带有噪声的'白银'标签作为替代。由于黄金标签仍是评估目标,直接聚合噪声标签可能引入偏差,而基于稀疏黄金标签的估计器方差较高,难以准确揭示模型性能差距。模型评估已从一次性任务演变为持续性的运营实践,跨越多个版本、发布周期和内容领域。一个自然问题是:能否利用过往历史评估数据改进当前评估?我们提出HERO(历史增强鲁棒评估)框架,通过历史数据抑制偏差(提升可靠性)并降低方差(提升敏感性)。HERO基于历史黄金标注学习白银标注者的性能,并通过高精度测量的历史协变量信息稳定估计结果。该方法可广泛应用于多种常见评估任务,即使仅部分历史标注者出现在当前轮次也有效。我们推导了偏差与方差降低的理论条件,在模拟研究中验证其性能,并在真实世界模型评估基准数据集上展示了其有效性。
原文摘要 · Abstract (English)
Reliable generative AI models critically rely on expert human annotations to evaluate output quality, yet these "gold" labels are expensive to collect and limited in quantity. Organizations thus often turn to collecting vast but noisy "silver" labels from crowdsourced workers or vendor annotators as proxies for gold labels. Because gold remains the evaluation target, naively aggregating noisy silver labels may introduce bias, and estimators built on sparsely observed gold labels may have high variance to resolve the model performance gaps that guide practical decisions. Model evaluation has become an ongoing operational practice rather than a one-time exercise, with evaluation rounds repeating across model versions, releases, and content domains. A natural question is whether the previous historical evaluation data can be used to improve each new round of evaluation. We introduce HERO (History Enhanced RObust model evaluation), a novel framework that uses historical data to suppress bias (improve reliability) and reduce variance (improve sensitivity) in model performance evaluation. HERO calibrates silver labelers' performance learned from historical gold annotations, and stabilizes the resulting estimator by anchoring it to covariate information measured with high precision in the historical data. HERO can be broadly applied across multiple common evaluation tasks, and remains valid when only a subset of historical labelers appears in the current round. We establish conditions under which the bias and variance reductions hold, showcase HERO's performance in simulation studies, and demonstrate its effectiveness on real-world model evaluation benchmarking datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。