解决标签噪声下置信集覆盖率不准问题,提升模型不确定性量化可靠性。
Noise-Adaptive Conformal Classification with Marginal Coverage
- 根据噪声水平自适应调整校准方式,缓解标签错误影响
- 在CIFAR-10H和BigEarthNet上实现接近理论最优的边际覆盖率
- 适合大规模含噪数据集上的可靠预测集生成任务
分位数推断为机器学习中的不确定性量化提供了严格的统计框架,可为任意分类模型生成校准良好的预测集,并保证精确的覆盖率。然而,其对理想数据可交换性的依赖限制了在真实世界复杂情况下的效果,例如低质量标签——这在现代大规模数据集中普遍存在。本文提出一种自适应分位数推断方法,能有效处理由随机标签噪声引起的可交换性偏离,在此类挑战性场景下仍可生成具有紧密边际覆盖率保证的信息性预测集。通过在合成数据和真实数据集(包括CIFAR-10H和BigEarthNet)上的广泛实验验证了该方法的有效性。
原文摘要 · Abstract (English)
Conformal inference provides a rigorous statistical framework for uncertainty quantification in machine learning, enabling well-calibrated prediction sets with precise coverage guarantees for any classification model. However, its reliance on the idealized assumption of perfect data exchangeability limits its effectiveness in the presence of real-world complications, such as low-quality labels -- a widespread issue in modern large-scale data sets. This work tackles this open problem by introducing an adaptive conformal inference method capable of efficiently handling deviations from exchangeability caused by random label noise, leading to informative prediction sets with tight marginal coverage guarantees even in those challenging scenarios. We validate our method through extensive numerical experiments demonstrating its effectiveness on synthetic and real data sets, including CIFAR-10H and BigEarthNet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。