arXiv:2606.08517cs.LGcs.CL2026-06被引 2

提出一种新方法,让模型在有限样本下安全地自适应选择预测,同时保证准确率和接受率。

A Joint Finite-Sample Certificate for Adaptive Selective Conformal Risk Control

论文配图:A Joint Finite-Sample Certificate for Adaptive Selective Conformal Risk Control
图 1 · 摘自论文原文
  • 直接建模风险比率,用改进的置信区间联合控制风险、接受率和效用
  • 在ImageNet和COCO上比现有方法提升22个百分点的接受率,且更紧致
  • 适合对安全性要求高的场景,如医疗或自动驾驶中的智能决策系统

选择性预测器只在高置信度输入上输出结果,其余情况则放弃预测;安全部署需一个有限样本证书,同时上界化选定风险、下界化接受概率(高于最小值 $\pmin$)及下界化部署效用。该证书需在从 $m$ 个阈值对组成的有限网格中自适应选择时依然有效,基于 $\ncert$ 个样本。本文针对有界但可能非单调的损失,不使用霍夫丁型区间估计,而是直接将选定风险视为比例,并结合三种置信界:方差自适应的经验伯恩斯坦界(风险)、克洛珀-皮尔逊界(接受率)、双向接近界(效用)。三者共同绝对下界化认证策略的效用,并在可行情况下,使其接近最优策略的效用至 $2\gammau$ 范围内。当风险裕度 $\gammar < \alpha$ 时,第三条腿与外部理想模型匹配,仅在此区域提供信息,而在主流操作点处不失效。相比仅依赖霍夫丁型比例构造,本方法将接受率下界依赖从 $1/\pmin$ 降至 $1/\sqrt{\pmin}$。闭式推论识别出每个阈值对的适用范围,在此范围内风险界优于霍夫丁-共形风险控制(Hoeffding--CRC)。实验显示,在ImageNet(三个ResNet)和COCO val 2017全景分割数据集上,证书使认证接受率提升22个百分点,且比非空泛基线紧约10倍;这些优势为特定场景所特有,不在ADE20K上出现。证书计算时间为 $O(\ncert m)$。

原文摘要 · Abstract (English)

Selective predictors answer on confident inputs and abstain elsewhere; deploying one safely needs a single finite-sample certificate that simultaneously upper-bounds the selected risk, lower-bounds the acceptance probability $\pacc$ above a floor $\pmin$, and lower-bounds the deployment utility. This certificate must be valid under adaptive threshold selection from a finite grid of $m$ pairs on $\ncert$ samples. We give such a certificate for bounded, possibly non-monotone losses by treating the selected risk directly as a ratio rather than through a Hoeffding-style range bound. The construction couples three confidence bounds: a variance-adaptive empirical-Bernstein bound on the ratio risk, a Clopper--Pearson bound on acceptance, and a two-sided closeness bound on utility. Together they lower-bound the certified policy's utility absolutely and to within $2\gammau$ of the best over the \emph{certified set}, both non-vacuous whenever feasible; a regime-scoped third leg matches an external oracle, informative only where the risk margin $\gammar < α$ and vacuous at the headline operating points. Relative to the range-only Hoeffding-ratio construction this sharpens the acceptance-floor dependence from $1/\pmin$ to $1/\sqrt{\pmin}$, and a closed-form corollary identifies a per-pair regime in which our risk bound dominates a Hoeffding conformal risk control (Hoeffding--CRC) selective bound. Empirically, on ImageNet (three ResNets) and COCO val 2017 panoptic, the certificate opens a $+22$ pp certified-acceptance frontier over Hoeffding--CRC and is ${\approx}10{\times}$ tighter than a non-vacuous matched-valid baseline; these gains are regime-scoped, not universal, and absent on ADE20K. The certifier runs in $O(\ncert m)$ time.

风险控制选择性预测置信界自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。