研究对抗攻击下分片置信预测的鲁棒性,可调控预测覆盖率。
Ensuring Calibration Robustness in Split Conformal Prediction Under Adversarial Attacks
- 用对抗扰动校准模型,实现覆盖概率可控
- 校准阶段引入攻击后,测试时覆盖率稳定在目标区间
- 训练阶段对抗训练可缩小预测集且保持信息量
置信预测(CP)提供无需分布假设、有限样本下的覆盖保证,但其依赖交换性条件,在分布偏移下常失效。本文研究测试时对抗扰动下分片置信预测的鲁棒性,关注覆盖有效性与预测集大小。理论分析揭示校准阶段对抗扰动强度对测试时覆盖保证的影响。进一步考察模型训练阶段的对抗训练效果。大量实验验证理论:(i) 预测覆盖率随校准阶段攻击强度单调变化,可通过设置非零校准攻击以可预测方式控制对抗测试下的覆盖率;(ii) 在合适校准攻击下,目标覆盖率可在一系列连续扰动水平内保持在指定容差带内;(iii) 训练阶段采用对抗训练可生成更紧凑的预测集,同时保持高信息性。
原文摘要 · Abstract (English)
Conformal prediction (CP) provides distribution-free, finite-sample coverage guarantees but critically relies on exchangeability, a condition often violated under distribution shift. We study the robustness of split conformal prediction under adversarial perturbations at test time, focusing on both coverage validity and the resulting prediction set size. Our theoretical analysis characterizes how the strength of adversarial perturbations during calibration affects coverage guarantees under adversarial test conditions. We further examine the impact of adversarial training at the model-training stage. Extensive experiments support our theory: (i) Prediction coverage varies monotonically with the calibration-time attack strength, enabling the use of nonzero calibration-time attack to predictably control coverage under adversarial tests; (ii) target coverage can hold over a range of test-time attacks: with a suitable calibration attack, coverage stays within any chosen tolerance band across a contiguous set of perturbation levels; and (iii) adversarial training at the training stage produces tighter prediction sets that retain high informativeness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。