提出可证明可靠的预测集,对抗数据投毒攻击。
Provably Reliable Conformal Prediction Sets in the Presence of Data Poisoning
- 用平滑得分函数聚合分块训练的分类器输出
- 多子集校准后取多数表决,提升抗投毒能力
- 在图像分类上既保准确又防攻击,适合安全场景
共形预测通过预测集提供模型无关且分布自由的不确定性量化,保证在任意用户指定概率下包含真实标签。然而,在攻击者同时篡改训练和校准数据的投毒攻击下,共形预测不可靠,实际预测集可能被严重扭曲。为此,我们提出可靠预测集(RPS):首个在投毒环境下具备可证明可靠性保障的高效方法。为应对训练投毒,引入平滑得分函数,可靠聚合基于不同训练子集训练的分类器输出;为应对校准投毒,构建多个基于不同校准子集的预测集,并通过多数表决生成最终预测集,仅当某类别出现在多数子集中才被包含。两种聚合策略均有效削弱了训练与校准数据中恶意样本的影响。我们在图像分类任务上实验验证,该方法在保持原始数据覆盖性的同时显著提升可靠性,整体推动了抗干扰不确定性量化的可信发展。
原文摘要 · Abstract (English)
Conformal prediction provides model-agnostic and distribution-free uncertainty quantification through prediction sets that are guaranteed to include the ground truth with any user-specified probability. Yet, conformal prediction is not reliable under poisoning attacks where adversaries manipulate both training and calibration data, which can significantly alter prediction sets in practice. As a solution, we propose reliable prediction sets (RPS): the first efficient method for constructing conformal prediction sets with provable reliability guarantees under poisoning. To ensure reliability under training poisoning, we introduce smoothed score functions that reliably aggregate predictions of classifiers trained on distinct partitions of the training data. To ensure reliability under calibration poisoning, we construct multiple prediction sets, each calibrated on distinct subsets of the calibration data. We then aggregate them into a majority prediction set, which includes a class only if it appears in a majority of the individual sets. Both proposed aggregations mitigate the influence of datapoints in the training and calibration data on the final prediction set. We experimentally validate our approach on image classification tasks, achieving strong reliability while maintaining utility and preserving coverage on clean data. Overall, our approach represents an important step towards more trustworthy uncertainty quantification in the presence of data poisoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。