修剪异常校准点能提升预测可靠性,但需满足特定条件。
When Does Trimming Help Conformal Prediction? A Retained-Law Diagnostic under Calibration Contamination
- 将修剪视为条件化而非净化,构建保留分布的理论框架。
- 发现修剪后覆盖率下降由干净数据与污染数据的保留比例决定。
- 提供可验证的有限样本保证,适合关注可信推理的研究者。
修剪可疑的校准点是应对共形预测中污染的常见做法。然而,其对干净目标覆盖率的影响由修剪引发的保留律决定,而非仅取决于污染程度。本文分析固定阈值修剪作为条件化而非净化的过程,将其替换为保留律,将覆盖率问题转化为一维得分-累积分布函数转移,并建立有限样本下的精确恒等式。通过分量边界控制转移差距,获得总体水平诊断:分离出干净侧协方差代价与保留污染代价,后者由脏数据到干净数据的保留比率决定。当异常得分能区分保留概率且对干净群体保持得分中立时,修剪才有效;否则无法通过保留混合系数显著降低污染。此外,还给出了独立审计下具有数值保障的有限样本证书模板。
原文摘要 · Abstract (English)
Trimming suspicious calibration points is a common response to contamination in conformal prediction. Its effect on clean-target coverage, however, is governed by the retained law induced by trimming, not by the contamination level alone. We analyse fixed-threshold trimming as conditioning rather than purification. It replaces the contaminated calibration law with a retained law, reducing clean-target coverage to a one-dimensional score-CDF transfer problem with an exact finite-sample identity. A componentwise bound on the transfer gap gives a population-level diagnostic. This separates a clean-side covariance cost from a retained-contamination cost, governed by the dirty-to-clean retention ratio. Trimming helps when the anomaly score separates retention probabilities while remaining score-neutral on the clean population. Otherwise, it cannot substantially reduce contamination through the retained mixture coefficient. We also give finite-sample certificate templates that provide numerical guarantees under independent audit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。