提出无需模型的图像攻击性度量方法,可直观理解攻击难易程度。
OTI: A Model-free and Visually Interpretable Measure of Image Attackability
- 基于语义物体纹理强度定义新度量OTI,不依赖模型梯度或扰动信息
- 在多个数据集上验证其与真实攻击成功率高度相关,计算效率高
- 结果可视化强,适合研究攻击机制和提升模型鲁棒性
尽管神经网络取得巨大成功,但良性图像仍可能被对抗扰动欺骗。有趣的是,不同图像的攻击性存在差异:在相同攻击条件下,某些图像易被破坏,而另一些则更抗干扰。评估图像攻击性在主动学习、对抗训练和攻击增强中具有重要意义。然而,现有方法稀少且存在两大局限:(1) 依赖模型代理提供先验信息(如梯度或最小扰动)来提取模型相关的图像特征;但实际中许多任务专用模型难以获取;(2) 提取的特征缺乏视觉可解释性,无法直接关联图像本身。为此,我们提出一种新型无模型、可视觉解释的图像攻击性度量——对象纹理强度(OTI),将图像攻击性量化为语义物体的纹理强度。理论上,从决策边界及对抗扰动的中高频特性两方面阐述了OTI原理。大量实验表明,OTI有效且计算高效,同时为对抗机器学习社区提供了对攻击性更直观的理解。
原文摘要 · Abstract (English)
Despite the tremendous success of neural networks, benign images can be corrupted by adversarial perturbations to deceive these models. Intriguingly, images differ in their attackability. Specifically, given an attack configuration, some images are easily corrupted, whereas others are more resistant. Evaluating image attackability has important applications in active learning, adversarial training, and attack enhancement. This prompts a growing interest in developing attackability measures. However, existing methods are scarce and suffer from two major limitations: (1) They rely on a model proxy to provide prior knowledge (e.g., gradients or minimal perturbation) to extract model-dependent image features. Unfortunately, in practice, many task-specific models are not readily accessible. (2) Extracted features characterizing image attackability lack visual interpretability, obscuring their direct relationship with the images. To address these, we propose a novel Object Texture Intensity (OTI), a model-free and visually interpretable measure of image attackability, which measures image attackability as the texture intensity of the image's semantic object. Theoretically, we describe the principles of OTI from the perspectives of decision boundaries as well as the mid- and high-frequency characteristics of adversarial perturbations. Comprehensive experiments demonstrate that OTI is effective and computationally efficient. In addition, our OTI provides the adversarial machine learning community with a visual understanding of attackability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。