提出新方法提升置信预测的自适应性,让难例输出更大预测集。
Quantifying and Improving Adaptivity in Conformal Prediction through Input Transformations
- 用输入变换排序样本难度,再均匀分组,解决传统分箱偏差。
- 新指标更准确评估覆盖误差与平均集合大小,相关性更强。
- 在图像分类和视力预测任务中表现优于现有方法,适合高可靠性场景。
置信预测通过输出标签集合而非单一预测值,同时提供概率覆盖保证。除覆盖保证外,对例子难度的自适应性也很重要:应为困难样本生成更大的预测集,简单样本则更小。现有自适应性评估方法常基于按难度分组的覆盖率违反或平均集合大小,但分组不平衡会导致覆盖或集合大小估计不准。为此,我们提出一种利用输入变换排序样本难度的分箱方法,并结合均匀质量分箱。基于此分箱,引入两个新指标以更可靠地评估自适应性。实验表明,新指标与期望自适应性相关性更强。进一步,受此发现启发,我们提出一种新的自适应预测集算法:按估计难度分组,对每组应用条件化置信预测,从而确定各组合适阈值。在(a)图像分类(ImageNet)和(b)医疗任务(视觉敏锐度预测)上的实验显示,该方法在新指标下优于现有方法。
原文摘要 · Abstract (English)
Conformal prediction constructs a set of labels instead of a single point prediction, while providing a probabilistic coverage guarantee. Beyond the coverage guarantee, adaptiveness to example difficulty is an important property. It means that the method should produce larger prediction sets for more difficult examples, and smaller ones for easier examples. Existing evaluation methods for adaptiveness typically analyze coverage rate violation or average set size across bins of examples grouped by difficulty. However, these approaches often suffer from imbalanced binning, which can lead to inaccurate estimates of coverage or set size. To address this issue, we propose a binning method that leverages input transformations to sort examples by difficulty, followed by uniform-mass binning. Building on this binning, we introduce two metrics to better evaluate adaptiveness. These metrics provide more reliable estimates of coverage rate violation and average set size due to balanced binning, leading to more accurate adaptivity assessment. Through experiments, we demonstrate that our proposed metric correlates more strongly with the desired adaptiveness property compared to existing ones. Furthermore, motivated by our findings, we propose a new adaptive prediction set algorithm that groups examples by estimated difficulty and applies group-conditional conformal prediction. This allows us to determine appropriate thresholds for each group. Experimental results on both (a) an Image Classification (ImageNet) (b) a medical task (visual acuity prediction) show that our method outperforms existing approaches according to the new metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。