为标签排序预测建立校准框架,揭示现有模型普遍不准。
Calibrated Preference Learning: The Case of Label Ranking
- 提出全排序、子排序、Top-k 排序的分级校准定义
- 实证发现主流模型在子排序与 Top-k 上校准度差异显著
- 校准度与基准准确率相关但不等价,反映新质量维度
校准指预测概率与实际发生频率的一致性,对可靠决策至关重要。尽管在分类和回归中已有广泛研究,标签排序的校准问题尚未被正式处理——其目标是预测标签集合顺序的分布。将排序简单视为类别会忽略其结构,无法捕捉成对关系和 Top-k 预测等关键特性。本文正式定义了标签排序的校准,并构建了涵盖全排序、子排序和 Top-k 排序的分层概念体系。证明全排序校准蕴含其他形式,但逆不成立;子排序与 Top-k 校准不可比较。实验表明,主流标签排序模型普遍校准不佳,子排序与 Top-k 指标间存在显著差异。应用于 RLHF 奖励模型时发现,校准度与基准准确率强相关但不完全一致,表明其捕获了超越 Top-1 准确率的重要质量维度。这些结果推动未来对校准偏差下游影响的理解及纠正方法的发展。
原文摘要 · Abstract (English)
Calibration, the alignment of predicted probabilities with true outcome frequencies, is essential for reliable decision-making. While extensively studied for classification and regression, calibration has not been formally addressed for probabilistic label ranking, where the goal is to predict a distribution over orderings of a label set. Naively treating rankings as classes ignores their structure and fails to capture important modalities such as pairwise and top-k predictions. We formalize calibration for label ranking and develop a hierarchy of notions covering full rankings, sub-rankings, and top-k rankings. We prove that full-rank calibration implies the others but not conversely, and sub-ranking and top-k calibration are incomparable. Empirically, we find popular label ranking models are often poorly calibrated, with substantial differences between sub-ranking and top-k metrics. Applying our framework to RLHF reward models, we find that calibration correlates strongly but not perfectly with benchmark accuracy, suggesting it captures a meaningful quality dimension beyond top-1 accuracy. These findings motivate future work on understanding the downstream effects of miscalibration and developing methods to correct it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。