arXiv:2602.16944cs.LG2026-02中稿 · the 23rd Internati…被引 2

用数学规划精确验证数据投毒攻击,确保防御无遗漏。

Exact Certification of Data-Poisoning Attacks Using Mixed-Integer Programming

  • 将训练与攻击建模为单一整数规划问题
  • 找到全局最优即得最坏情况投毒效果
  • 首次实现训练阶段鲁棒性的严格证明

本文提出一种验证框架,可对神经网络训练过程中的数据投毒攻击提供可靠且完整的保证。将对抗性数据操纵、模型训练和测试评估统一建模为一个混合整数二次规划(MIQCP)问题。求解该模型的全局最优,可严格确定最坏情况下的投毒攻击效果,同时对所有可能攻击在给定训练流程下的有效性进行上界约束。该框架同时编码了基于梯度的训练动态和测试时的模型评估,实现了训练阶段鲁棒性的首次精确认证。在小型模型上的实验表明,本方法能够完整刻画对数据投毒的抗性能力。

原文摘要 · Abstract (English)

This work introduces a verification framework that provides both sound and complete guarantees for data poisoning attacks during neural network training. We formulate adversarial data manipulation, model training, and test-time evaluation in a single mixed-integer quadratic programming (MIQCP) problem. Finding the global optimum of the proposed formulation provably yields worst-case poisoning attacks, while simultaneously bounding the effectiveness of all possible attacks on the given training pipeline. Our framework encodes both the gradient-based training dynamics and model evaluation at test time, enabling the first exact certification of training-time robustness. Experimental evaluation on small models confirms that our approach delivers a complete characterization of robustness against data poisoning.

数据投毒安全验证整数规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。