arXiv:2509.05826cs.LGcs.CV2025-09中稿 · the IEEE/CVF Winte…被引 1

检验了置信预测在捕捉数据固有不确定性上的表现

Performance of Conformal Prediction in Capturing Aleatoric Uncertainty

  • 用人类标注者对同一实例的多标签一致性衡量不确定性
  • 多数情况下预测集大小与人类标注相关性很弱或弱
  • 适合关注预测可靠性但需警惕其对不确定性的误判

置信预测是一种模型无关的方法,可生成以高概率包含真实类别的预测集合。尽管其预测集大小理论上应反映数据的随机不确定性,但缺乏实证支持。现有研究认为预测集大小可上界随机不确定性,或对困难样本生成更大集合,但这一属性尚未验证。本文通过测量预测集大小与人类标注者对同一实例所分配的不同标签数量之间的相关性,评估置信预测量化随机不确定性的能力。同时比较预测集与人类标注的一致性。实验使用三种置信预测方法,在四个数据集上对八种深度学习模型生成预测集,这些数据集每个实例均有5至50名标注者参与,便于识别类别重叠。结果显示,绝大多数置信预测输出与人类标注的相关性极弱或弱,仅有少数显示中等相关性。这表明置信预测虽能提高真实类别的覆盖概率,但在捕捉随机不确定性及生成与人类标注一致的预测集方面仍存在局限,亟需重新审视其生成的预测集。

原文摘要 · Abstract (English)

Conformal prediction is a model-agnostic approach to generating prediction sets that cover the true class with a high probability. Although its prediction set size is expected to capture aleatoric uncertainty, there is a lack of evidence regarding its effectiveness. The literature presents that prediction set size can upper-bound aleatoric uncertainty or that prediction sets are larger for difficult instances and smaller for easy ones, but a validation of this attribute of conformal predictors is missing. This work investigates how effectively conformal predictors quantify aleatoric uncertainty, specifically the inherent ambiguity in datasets caused by overlapping classes. We perform this by measuring the correlation between prediction set sizes and the number of distinct labels assigned by human annotators per instance. We further assess the similarity between prediction sets and human-provided annotations. We use three conformal prediction approaches to generate prediction sets for eight deep learning models trained on four datasets. The datasets contain annotations from multiple human annotators (ranging from five to fifty participants) per instance, enabling the identification of class overlap. We show that the vast majority of the conformal prediction outputs show a very weak to weak correlation with human annotations, with only a few showing moderate correlation. These findings underscore the necessity of critically reassessing the prediction sets generated using conformal predictors. While they can provide a higher coverage of the true classes, their capability in capturing aleatoric uncertainty and generating sets that align with human annotations remains limited.

置信预测不确定性机器学习标注一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。