arXiv:2510.15699cs.LG2025-10

让对抗样本符合真实数据约束,提升攻击真实性和效率

Constrained Adversarial Perturbation

  • 基于增强拉格朗日法构建带约束的对抗攻击框架
  • 在金融、网络等多领域实现更高成功率与更快运行速度
  • 可自动学习数据约束,适用于结构化输入场景

深度神经网络在众多分类任务中取得显著成果,但仍易受对抗样本影响——这些样本经微小扰动后导致误分类,而人类难以察觉。在各类攻击策略中,通用对抗扰动(UAP)成为测试模型鲁棒性及支持可扩展对抗训练的强大工具。然而,现有大多数UAP方法忽略了制约特征关系的领域特定约束。违反此类约束(如信贷评分中的债务收入比或网络通信中的包流不变性)会使对抗样本变得不切实际或易于检测,限制其实际应用。本文提出受限对抗扰动(CAP),通过构建基于增强拉格朗日的最小最大优化问题,施加多种可能复杂的约束,且可处理不同重要性层级。我们设计了一种基于梯度的交替优化算法,高效求解该问题。在金融、网络通信和网络物理系统等多个领域评估显示,CAP在保持更高攻击成功率的同时显著降低运行时间。该方法还可无缝推广至个体对抗扰动,同样表现出优异性能。最后,我们提出一种从数据中学习特征约束的系统性方法,使本框架在具有结构化输入空间的各类场景中具备广泛适用性。

原文摘要 · Abstract (English)

Deep neural networks have achieved remarkable success in a wide range of classification tasks. However, they remain highly susceptible to adversarial examples - inputs that are subtly perturbed to induce misclassification while appearing unchanged to humans. Among various attack strategies, Universal Adversarial Perturbations (UAPs) have emerged as a powerful tool for both stress testing model robustness and facilitating scalable adversarial training. Despite their effectiveness, most existing UAP methods neglect domain specific constraints that govern feature relationships. Violating such constraints, such as debt to income ratios in credit scoring or packet flow invariants in network communication, can render adversarial examples implausible or easily detectable, thereby limiting their real world applicability. In this work, we advance universal adversarial attacks to constrained feature spaces by formulating an augmented Lagrangian based min max optimization problem that enforces multiple, potentially complex constraints of varying importance. We propose Constrained Adversarial Perturbation (CAP), an efficient algorithm that solves this problem using a gradient based alternating optimization strategy. We evaluate CAP across diverse domains including finance, IT networks, and cyber physical systems, and demonstrate that it achieves higher attack success rates while significantly reducing runtime compared to existing baselines. Our approach also generalizes seamlessly to individual adversarial perturbations, where we observe similar strong performance gains. Finally, we introduce a principled procedure for learning feature constraints directly from data, enabling broad applicability across domains with structured input spaces.

对抗攻击约束优化数据可信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。