arXiv:2412.10186cs.LGcs.AI2024-12NeurIPS被引 1

用数学方法证明模型在数据被篡改时仍可靠,防攻击更安心。

MIBP-Cert: Certified Training against Data Perturbations with Mixed-Integer Bilinear Programs

  • 基于混合整数双线性规划,严格计算模型对数据扰动的鲁棒边界。
  • 可应对连续与离散数据,覆盖复杂攻击场景,给出确定性安全保证。
  • 适合关注模型安全性、对抗训练或可信AI的研究者与开发者。

训练过程中数据错误、损坏和污染攻击严重威胁现代AI系统的可靠性。尽管已有大量经验性缓解措施,但攻击不断演化且数据复杂,亟需一种更严谨、可证明的方法来应对这些挑战,并理解扰动如何影响最终模型。为此,我们提出MIBP-Cert,一种基于混合整数双线性规划(MIBP)的新认证方法,能计算出可靠且确定的边界,在复杂威胁模型下提供可证明的鲁棒性。通过计算受扰动或被操纵数据可达的参数集合,可预测所有可能结果并确保模型稳健。为使该优化问题可解,我们设计了一种新松弛方案,在不牺牲保真性的前提下约束每一步训练。实验表明,该方法适用于连续与离散数据,以及多种威胁模型——包括此前难以处理的复杂情形。

原文摘要 · Abstract (English)

Data errors, corruptions, and poisoning attacks during training pose a major threat to the reliability of modern AI systems. While extensive effort has gone into empirical mitigations, the evolving nature of attacks and the complexity of data require a more principled, provable approach to robustly learn on such data - and to understand how perturbations influence the final model. Hence, we introduce MIBP-Cert, a novel certification method based on mixed-integer bilinear programming (MIBP) that computes sound, deterministic bounds to provide provable robustness even under complex threat models. By computing the set of parameters reachable through perturbed or manipulated data, we can predict all possible outcomes and guarantee robustness. To make solving this optimization problem tractable, we propose a novel relaxation scheme that bounds each training step without sacrificing soundness. We demonstrate the applicability of our approach to continuous and discrete data, as well as different threat models - including complex ones that were previously out of reach.

鲁棒学习形式化验证对抗攻击优化建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。