arXiv:2507.07776cs.CV2025-07

构建可量化评估无约束对抗样本真实感的开源框架

SCOOTER: A Human Evaluation Framework for Unrestricted Adversarial Examples

  • 提出基于众包的统计评估规范,确保人类评价可靠
  • 346人测试发现六类攻击均无法生成难以察觉的干扰图像
  • 提供浏览器工具与数据集,适合安全研究者和评测人员使用

无约束对抗攻击突破传统范数限制,可通过改变物体颜色等方式欺骗视觉模型,绕过常规防御机制。但由于缺乏感知不可见性保证,需通过人工评估验证其真实性。现有工作缺乏统计效力,难以系统比较。为此,我们提出SCOOTER——首个面向无约束对抗样本的开源、统计驱动评估框架。贡献包括:(i) 提出众包实验中的功率、补偿与李克特量表等效性标准;(ii) 首次大规模人类-模型对比(346名参与者)显示三种颜色空间攻击和三种扩散攻击均无法生成不可察觉的对抗样本;且GPT-4o仅对其中四种攻击能稳定检测;(iii) 开源浏览器任务模板及Python/R分析脚本;(iv) 构建基于ImageNet的数据集,含3000张真实图像、7000个对抗样本和超过3.4万条人工评分。结果表明自动视觉系统与人类感知不一致,亟需以SCOOTER为基准的真实评估。

原文摘要 · Abstract (English)

Unrestricted adversarial attacks aim to fool computer vision models without being constrained by $\ell_p$-norm bounds to remain imperceptible to humans, for example, by changing an object's color. This allows attackers to circumvent traditional, norm-bounded defense strategies such as adversarial training or certified defense strategies. However, due to their unrestricted nature, there are also no guarantees of norm-based imperceptibility, necessitating human evaluations to verify just how authentic these adversarial examples look. While some related work assesses this vital quality of adversarial attacks, none provide statistically significant insights. This issue necessitates a unified framework that supports and streamlines such an assessment for evaluating and comparing unrestricted attacks. To close this gap, we introduce SCOOTER - an open-source, statistically powered framework for evaluating unrestricted adversarial examples. Our contributions are: $(i)$ best-practice guidelines for crowd-study power, compensation, and Likert equivalence bounds to measure imperceptibility; $(ii)$ the first large-scale human vs. model comparison across 346 human participants showing that three color-space attacks and three diffusion-based attacks fail to produce imperceptible images. Furthermore, we found that GPT-4o can serve as a preliminary test for imperceptibility, but it only consistently detects adversarial examples for four out of six tested attacks; $(iii)$ open-source software tools, including a browser-based task template to collect annotations and analysis scripts in Python and R; $(iv)$ an ImageNet-derived benchmark dataset containing 3K real images, 7K adversarial examples, and over 34K human ratings. Our findings demonstrate that automated vision systems do not align with human perception, reinforcing the need for a ground-truth SCOOTER benchmark.

对抗样本人类评估计算机视觉基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。