用少量人工评分+大量自动评分,高效且无偏地比较模型性能
Which Metrics Save the Most Human Annotation? Prediction-Powered Evaluation and Meta-Evaluation

- 结合少量人工打分与大规模自动评分,实现数据高效评估
- 在6个WMT数据集上验证,结果无偏且显著降低人工成本
- 提出新元指标PPSR,可量化自动指标节省的人工量
在多种不可验证任务中,人工评估可靠但成本高,自动指标可扩展但常有偏差。基于预测驱动推理(PPI),本文提出预测驱动评估框架,将有限人工判断与大规模自动评分结合,实现数据高效、理论上无偏的系统比较。我们设计了参数化与非参数化方法,分析成对与非成对设计间的效率权衡,并在六个WMT数据集上验证该框架。进一步提出预测驱动节省率(PPSR),一种衡量自动指标在预测驱动评估中能节省多少人工标注的元指标。PPSR直接针对预测驱动评估中的指标效用,比现有系统级元指标更具区分性与稳定性。总体而言,本研究将自动指标重新定位为降低人工标注成本的工具,而非替代人工判断,适用于广泛不可验证任务。
原文摘要 · Abstract (English)
Across various non-verifiable tasks, human evaluation is reliable but expensive, while automatic metrics are more scalable but often biased. Building on prediction-powered inference (PPI), we propose prediction-powered evaluation, a framework that combines limited human judgments with large-scale automatic scores to obtain data-efficient system comparisons that are provably unbiased. We develop parametric and non-parametric procedures, analyze the efficiency trade-off between paired and unpaired designs, and validate the framework on six WMT datasets. We further introduce the Prediction-Powered Saving Ratio (PPSR), a meta-metric that measures how much human annotation an automatic metric can save when used within prediction-powered evaluation. PPSR directly targets metric utility for prediction-powered evaluation and yields more discriminative and stable metric rankings than existing system-level meta-metrics. Overall, our new paradigm reframes automatic metrics as tools for reducing human annotation cost rather than replacing human judgment, and applies broadly to non-verifiable tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。