arXiv:2505.04986stat.MLcs.LG2025-05ICML被引 4

针对测试数据中的异常值,提出先检测后填补的校准方法,提升预测区间可靠性。

Conformal Prediction with Cellwise Outliers: A Detect-then-Impute Approach

  • 先检测测试特征中的异常值,再用插补法修复,恢复数据交换性
  • 在合成与真实数据上验证,覆盖率达1-2α且效率接近理想基准
  • 适合处理含异常值的黑箱模型预测,尤其关注鲁棒性场景

conformal prediction 能为黑箱模型构建具有有限样本覆盖率保证的预测区间,但当测试特征中存在异常值(如单元格异常)时,其交换性假设被破坏。本文提出一种新的“先检测后填补”框架:首先对测试特征进行异常检测,再用插补方法修复被识别为异常的单元格。为量化处理后特征的不确定性,我们自适应地将检测与插补过程应用于校准集,从而构建可用于测试标签预测区间的交换性特征。开发了两个实用算法 PDI-CP 与 JDI-CP,证明在常用检测与插补方法下具有分布无关的覆盖率分析,其中 JDI-CP 实现有限样本 $1-2α$ 的覆盖率保证。在合成与真实数据集上的实验表明,所提算法具备稳健的覆盖率表现,且计算效率接近理想基准。

原文摘要 · Abstract (English)

Conformal prediction is a powerful tool for constructing prediction intervals for black-box models, providing a finite sample coverage guarantee for exchangeable data. However, this exchangeability is compromised when some entries of the test feature are contaminated, such as in the case of cellwise outliers. To address this issue, this paper introduces a novel framework called detect-then-impute conformal prediction. This framework first employs an outlier detection procedure on the test feature and then utilizes an imputation method to fill in those cells identified as outliers. To quantify the uncertainty in the processed test feature, we adaptively apply the detection and imputation procedures to the calibration set, thereby constructing exchangeable features for the conformal prediction interval of the test label. We develop two practical algorithms, PDI-CP and JDI-CP, and provide a distribution-free coverage analysis under some commonly used detection and imputation procedures. Notably, JDI-CP achieves a finite sample $1-2α$ coverage guarantee. Numerical experiments on both synthetic and real datasets demonstrate that our proposed algorithms exhibit robust coverage properties and comparable efficiency to the oracle baseline.

预测区间异常检测插补鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。