提出统一框架,可证明模型对训练数据篡改的鲁棒性。
Abstract Gradient Training: A Unified Certification Framework for Data Poisoning, Unlearning, and Differential Privacy
- 通过参数空间边界分析,统一处理数据污染、删除和隐私保护问题。
- 首次为梯度优化训练的模型提供可证明的训练数据扰动鲁棒性保证。
- 适合关注模型安全、数据隐私和可撤销学习的研究者使用。
推理阶段的数据扰动(如对抗攻击)在机器学习中已得到广泛研究,相关鲁棒性认证技术已较为成熟。相比之下,针对训练数据扰动的认证研究仍相对不足。这类扰动主要出现在三种关键场景:对抗性数据污染(攻击者操纵训练样本以破坏模型性能)、机器可撤销学习(需验证移除特定训练数据后模型行为的稳定性)以及差分隐私(需对替换单个数据点提供保障)。本文提出抽象梯度训练(AGT),一种统一的认证框架,用于评估给定模型及训练过程对训练数据扰动的鲁棒性,涵盖有界扰动、数据点移除及新样本添加。通过界定参数可达集(即建立可证明的参数空间边界),AGT为基于一阶优化方法训练的模型提供形式化分析手段。
原文摘要 · Abstract (English)
The impact of inference-time data perturbation (e.g., adversarial attacks) has been extensively studied in machine learning, leading to well-established certification techniques for adversarial robustness. In contrast, certifying models against training data perturbations remains a relatively under-explored area. These perturbations can arise in three critical contexts: adversarial data poisoning, where an adversary manipulates training samples to corrupt model performance; machine unlearning, which requires certifying model behavior under the removal of specific training data; and differential privacy, where guarantees must be given with respect to substituting individual data points. This work introduces Abstract Gradient Training (AGT), a unified framework for certifying robustness of a given model and training procedure to training data perturbations, including bounded perturbations, the removal of data points, and the addition of new samples. By bounding the reachable set of parameters, i.e., establishing provable parameter-space bounds, AGT provides a formal approach to analyzing the behavior of models trained via first-order optimization methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。