arXiv:2506.05382cs.CRcs.AI2025-06

提出新攻击方法,让对抗样本更难被发现且更鲁棒。

How stealthy is stealthy? Studying the Efficacy of Black-Box Adversarial Attacks in the Real World

  • 用高斯模糊梯度采样+局部代理模型生成攻击
  • 在压缩、自动检测、人工观察下均表现更隐蔽
  • 适合研究真实场景下模型安全性的研究人员

深度学习系统在自动驾驶等关键领域易受对抗样本影响。本研究聚焦计算机视觉中的黑盒对抗攻击,即攻击者仅能查询目标模型。引入三个评估指标:对压缩的鲁棒性、自动检测下的隐蔽性、人工观察下的隐蔽性。现有方法往往在三者间权衡不足。本文提出ECLIPSE方法,采用采样梯度上的高斯模糊与局部代理模型。在公开数据集上的全面实验表明,ECLIPSE在三者间取得更好平衡,显著提升了攻击的实用性。

原文摘要 · Abstract (English)

Deep learning systems, critical in domains like autonomous vehicles, are vulnerable to adversarial examples (crafted inputs designed to mislead classifiers). This study investigates black-box adversarial attacks in computer vision. This is a realistic scenario, where attackers have query-only access to the target model. Three properties are introduced to evaluate attack feasibility: robustness to compression, stealthiness to automatic detection, and stealthiness to human inspection. State-of-the-Art methods tend to prioritize one criterion at the expense of others. We propose ECLIPSE, a novel attack method employing Gaussian blurring on sampled gradients and a local surrogate model. Comprehensive experiments on a public dataset highlight ECLIPSE's advantages, demonstrating its contribution to the trade-off between the three properties.

对抗攻击黑盒攻击模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。