揭示黑盒攻击中隐藏的先验信息偏差,提出无先验的公平评估框架。
On Transfer-based Universal Attacks in Pure Black-box Setting
- 构建无先验假设的透明评估框架,消除数据集和类别数等隐含依赖。
- 发现现有方法因依赖先验导致迁移性评分虚高,真实性能被严重夸大。
- 提出图像融合新策略,提升代理模型训练效果,适用于实际查询攻击场景。
尽管深度视觉模型表现优异,却易受可迁移黑盒对抗攻击影响。现有方法通常以目标模型无关方式生成扰动,但研究发现其无意中依赖了违反黑盒假设的各类先验信息,如目标模型训练数据集及类别数量知识。这使得该领域对可迁移黑盒攻击的真实威力缺乏准确评估。本文通过实证研究揭示这些偏差,并提出一个无先验的透明分析框架。基于此框架,我们分析了目标模型数据与类别数先验对攻击性能的影响,发现这些先验导致迁移性分数被显著高估。此外,我们将框架扩展至查询型攻击,提出一种新颖的图像融合技术,用于高效构建代理模型训练数据。
原文摘要 · Abstract (English)
Despite their impressive performance, deep visual models are susceptible to transferable black-box adversarial attacks. Principally, these attacks craft perturbations in a target model-agnostic manner. However, surprisingly, we find that existing methods in this domain inadvertently take help from various priors that violate the black-box assumption such as the availability of the dataset used to train the target model, and the knowledge of the number of classes in the target model. Consequently, the literature fails to articulate the true potency of transferable black-box attacks. We provide an empirical study of these biases and propose a framework that aids in a prior-free transparent study of this paradigm. Using our framework, we analyze the role of prior knowledge of the target model data and number of classes in attack performance. We also provide several interesting insights based on our analysis, and demonstrate that priors cause overestimation in transferability scores. Finally, we extend our framework to query-based attacks. This extension inspires a novel image-blending technique to prepare data for effective surrogate model training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。