提升搜索结果可靠性排序,让可信度更精准匹配真实相关性。
Post-Calibration Reliability Reranking of Relevance Decisions via Label-wise Monotone Projection

- 基于标签级单调投影,重新校准置信度以反映真实可靠性。
- 在6个数据集上显著提升可靠性排序效果与回退决策效率。
- 适合关注搜索系统可信度评估与优化的工程师与研究者。
网页搜索、商品搜索和问答检索系统通常为每个查询-候选对分配相关性标签和置信度分数。置信度常用于指导下游决策或备选方案选择。由于置信度与实际正确率不一致,需进行事后校准。然而,现有校准方法仅使置信度与平均正确率对齐,未能消除同一置信水平内不同预测标签间的可靠性差异。为此,本文提出标签级单调可靠性投影(MRP),学习标签相关的单调函数,将校准后的置信度映射为正确性可靠性,同时保留原始预测标签和类别概率。该方法通过残差风险重排固定预测结果,提升可靠性排序质量。在六个信息检索相关性数据集及多种事后校准器上,MRP均显著改善可靠性排序与平均回退效用,且保持全覆盖准确率和ECE。结构消融分析表明,主要收益来自标签级残差可靠性,而非全局置信度重映射。进一步分析显示,该投影可用于顶层标签概率空间的兼容性分析,但其本质区别于核心可靠性重排序目标。代码将公开发布。
原文摘要 · Abstract (English)
Web search, product search, and question-answering retrieval systems often assign a relevance label and confidence score to each query-candidate pair. The relevance label describes how well a page, product, or passage matches the query, while the confidence often guides downstream use or fallback decisions. Post-hoc calibration is therefore needed because misaligned confidence can make systems over-trust wrong predictions or unnecessarily defer correct ones. However, calibration mainly aligns confidence with average correctness, and does not remove predicted-label-dependent reliability differences that remain within the same calibrated confidence level. We address this gap with Label-wise Monotone Reliability Projection (MRP), which learns label-wise monotone functions that map calibrated confidence to correctness reliability while preserving the original predicted labels and class probabilities. The resulting reliability score reranks fixed predictions according to residual risk. Across six information access relevance datasets and multiple post-hoc calibrators, MRP improves reliability reranking and average fallback utility while preserving full-coverage accuracy and ECE. Structural ablations show that the main gains come from label-wise residual reliability rather than from global confidence remapping. We further analyze when MRP reliability scores can be embedded back into top-label probability geometry, showing that this projection is useful as a compatibility analysis but is distinct from the main reliability-reranking objective. The implementation will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。