用少量人工标注校准多个AI评判,提升模型评估的准确性与可靠性。
BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges

- 结合少量人工标注与多AI评判结果,构建校准模型。
- 在不同任务中降低偏差与方差,提升排名一致性。
- 适合资源有限时进行高可靠性的模型评估。
AI评判提供了可扩展、低成本的人类评估替代方案,但其输出可能偏离人类偏好且高度依赖具体项目,不同评判者、任务和领域间差异显著。若直接使用未经校准的AI评价进行模型排序、项目评分或群体质量报告,会直接影响下游决策。本文提出BACON,一种四阶段流程,融合预算内人工校准与多AI评判输出,生成更准确的标注。BACON为每项内容构建全覆盖辅助特征,包括多评判得分、分词级别不确定性统计与上下文嵌入。随后对小样本进行人工标注,并训练交叉拟合结果模型,生成校准后的项目级代理预测。这些预测支持两类应用:基于增强估计方程的群体级汇总指标(如均值、分位数)估计,附带有效置信区间;以及个体级代理打分,用于项目排序与标注。BACON将AI评判视为辅助测量而非真实标签:人工标注提供校准锚点,而AI信号提升效率。在多样任务、领域与标注预算下,BACON均优于原始AI输出与纯人工方法,在预测准确率、排名一致性方面表现更优,同时降低偏差与方差。结果表明,BACON为有限人工标注条件下的可扩展评估提供了实用且统计严谨的框架。
原文摘要 · Abstract (English)
AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges, tasks, and domains. When uncalibrated AI evaluations are used for model ranking, item scoring, or population-level quality reporting, these biases can directly distort downstream decisions. We propose BACON, a four-stage pipeline that combines budgeted human calibration with multiple AI-judge outputs to produce more accurate annotations. BACON constructs full-coverage auxiliary features for every item, including multi-judge scores, token-level uncertainty statistics, and contextual embeddings. It then collects human labels for a small sampled subset and trains a cross-fitted outcome model to generate calibrated item-level surrogate predictions. These predictions support two use cases: population-level estimation of summary metrics, such as means or quantiles, using an augmented estimating-equation estimator with valid confidence intervals; and individual-level surrogate scoring for item ranking and annotation. BACON treats AI judges as auxiliary measurements rather than ground truth: human labels provide the calibration anchor, while AI-derived signals improve efficiency. Across diverse tasks, domains, and labeling budgets, BACON improves predictive accuracy and ranking consistency, and reduces bias and variance relative to raw AI outputs and purely human-label-based methods. These results show that BACON offers a practical, statistically grounded framework for scalable evaluation with limited human annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。