arXiv:2608.02455cs.LGstat.ML2026-08中稿 · ICLR

融合人类判断与模型评分,提升无真值场景下的评估准确性

Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees

论文配图:Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees
图 1 · 摘自论文原文
  • 先聚合不同专家的对比判断生成共识排序
  • 再通过保序投影校准模型分数,提升一致性
  • 理论保证在数据不全时仍有效,适合无真值评估

人类中心评估任务依赖人工判断,通常缺乏可验证的真值。现有方法面临两难:仅用人类判断易受评者能力差异和评分尺度不一致影响;仅用模型评分则需依赖不完美的代理标签或不完整特征。本文提出「先聚合后校准」(AtC)框架:第一阶段利用考虑标注者可靠性的排名聚合模型,将异质性比较判断整合为共识排序;第二阶段通过保序投影将任意预测模型的得分校准至该顺序,既保证序数一致性,又尽可能保留模型的定量信息。理论上证明:(1) 考虑标注者异质性可实现更高效的共识估计;(2) 保序校准即使在共识排序错误时也具备风险界;(3) AtC渐近优于纯模型评估。在半合成与真实数据集上,AtC持续提升准确率与鲁棒性。结果连接了判断聚合与模型无关校准,为真值不可靠或稀缺场景提供可信赖的人类中心评估方案。

原文摘要 · Abstract (English)

Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth. Existing approaches face a dilemma: methods using only human judgments suffer from heterogeneous expertise and inconsistent rating scales, while methods using only model-generated scores must learn from imperfect proxies or incomplete features. We propose Aggregate-then-Calibrate (AtC), a two-stage framework that combines these complementary sources. Stage-1 aggregates heterogeneous comparative judgments into a consensus ranking using a rank-aggregation model that accounts for annotator reliability. Stage-2 calibrates any predictive model's scores by an isotonic projection onto the order, enforcing ordinal consistency while preserving as much of the model's quantitative information as possible. Theoretically, we show: (1) modeling annotator heterogeneity yields strictly more efficient consensus estimation than homogeneity; (2) isotonic calibration enjoys risk bounds even when the consensus ranking is misspecified; and (3) AtC asymptotically outperforms model-only assessment. Across semi-synthetic and real-world datasets, AtC consistently improves accuracy and robustness over human-only or model-only assessments. Our results bridge judgment aggregation with model-free calibration, providing a principled recipe for human-centered assessment when ground truth is costly, scarce, or unverifiable.

评估系统人类判断保序校准共识排序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。