提出可生成结构化稀疏扰动的攻击框架,提升模型解释性。
A Versatile Framework for Designing Group-Sparse Adversarial Attacks
- 基于重叠L0正则化,实现像素与通道分组的结构化扰动优化。
- 在CIFAR-10和ImageNet上实现100%攻击成功率,扰动更稀疏且结构清晰。
- 适用于白盒攻击,可定位关键特征区域,提供反事实解释。
现有对抗攻击常忽略扰动稀疏性,难以建模结构变化并解释深度神经网络对有意义输入模式的处理机制。本文提出ATOS(Attack Through Overlapping Sparsity)——一种可微优化框架,支持逐元素、逐像素及分组形式的结构化稀疏对抗扰动生成。针对图像分类器的白盒攻击,引入重叠平滑L0(OSL0)函数,在促进收敛至驻点的同时,鼓励稀疏且结构化的扰动。通过分组通道与邻近像素,ATOS增强可解释性,有助于识别鲁棒与非鲁棒特征。采用对数-指数和近似L-infinity梯度,精确控制扰动幅度。在CIFAR-10与ImageNet上,ATOS实现100%攻击成功率,生成的扰动显著更稀疏且结构更连贯。分组攻击可突出网络视角下的关键区域,通过替换类别定义区域为目标类鲁棒特征,提供反事实解释。
原文摘要 · Abstract (English)
Existing adversarial attacks often neglect perturbation sparsity, limiting their ability to model structural changes and to explain how deep neural networks (DNNs) process meaningful input patterns. We propose ATOS (Attack Through Overlapping Sparsity), a differentiable optimization framework that generates structured, sparse adversarial perturbations in element-wise, pixel-wise, and group-wise forms. For white-box attacks on image classifiers, we introduce the Overlapping Smoothed L0 (OSL0) function, which promotes convergence to a stationary point while encouraging sparse, structured perturbations. By grouping channels and adjacent pixels, ATOS improves interpretability and helps identify robust versus non-robust features. We approximate the L-infinity gradient using the logarithm of the sum of exponential absolute values to tightly control perturbation magnitude. On CIFAR-10 and ImageNet, ATOS achieves a 100% attack success rate while producing significantly sparser and more structurally coherent perturbations than prior methods. The structured group-wise attack highlights critical regions from the network's perspective, providing counterfactual explanations by replacing class-defining regions with robust features from the target class.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。