分离优化提升预测集效率,兼顾准确率与紧凑性。
Decoupled Conformal Optimisation: Efficient Prediction Sets via Independent Tuning and Calibration

- 用独立数据集分别优化结构和校准阈值,避免传统方法的耦合瓶颈。
- 在多个数据集上平均预测集大小减少,95%分位数也显著缩小。
- 适合追求高效且可靠预测集的研究者或工业落地场景。
贝叶斯置信校准方法通常使用相同预留数据同时搜索高效的预测集并校验覆盖概率或风险。这种耦合对高概率风险控制是自然的,但在目标为标准有限样本边际覆盖时并非必要。本文提出解耦置信优化(DCO),采用训练-调优-校准的设计范式:用独立调优集进行结构选择以提升效率,再用全新校准集确定最终的置信阈值。在给定调优结构条件下,基于分割置信交换性,可保证任意候选类别具有有限样本边际覆盖,无需置信参数或多测试校正。因此,DCO针对的是不同于PAC类方法的有限样本保证:边际置信覆盖而非高概率风险控制。在一致假设下,两种方法收敛于相同的总体阈值。在分类与回归基准(包括ImageNet-A、CIFAR-100、Diabetes、California Housing和Concrete)上,DCO能紧密跟踪名义覆盖水平,同时显著降低平均预测集大小或区间宽度。例如在ImageNet-A上,平均集大小从26.52降至25.26,第95百分位从58.95降至53.73;在Diabetes数据集上,平均区间宽度由2.098降至1.914。
原文摘要 · Abstract (English)
Bayesian conformal optimisation methods often use the same held-out data both to search for efficient prediction sets and to certify coverage or risk. This coupling is natural for high-probability risk-control guarantees, but it is not necessary when the target is standard finite-sample marginal conformal coverage. We propose Decoupled Conformal Optimisation (DCO), a train-tune-calibrate design principle that uses an independent tuning split for efficiency-oriented structural selection and a fresh calibration split for the final conformal quantile. Conditional on the tuned structure, standard split-conformal exchangeability yields finite-sample marginal coverage for any candidate class, without a confidence parameter or multiple-testing correction. DCO therefore targets a different finite-sample guarantee from PAC-style methods: marginal conformal coverage rather than high-probability risk control. Under consistency assumptions on the coupled risk bound, the two approaches nevertheless converge to the same population threshold. Across classification and regression benchmarks, including ImageNet-A, CIFAR-100, Diabetes, California Housing, and Concrete, DCO tracks the nominal coverage level closely while often reducing average prediction-set size or interval width relative to PAC-style calibration. On ImageNet-A, for example, the average set size decreases from $26.52$ to $25.26$ and the 95th-percentile set size from $58.95$ to $53.73$; on Diabetes, the average interval width decreases from $2.098$ to $1.914$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。