arXiv:2409.05558cs.CVcs.AI2024-09被引 4

通过添加不同强度的遮罩,可有效欺骗主流图像分类模型。

Seeing Through the Mask: Rethinking Adversarial Examples for CAPTCHAs

  • 用可感知但不影响人类识别的遮罩干扰图像
  • 所有模型Acc@1下降超50个百分点,视觉变压器下降80个百分点
  • 揭示当前机器视觉仍难超越人类认知能力

现代CAPTCHA严重依赖计算机难以处理而人类轻松应对的视觉任务。然而,图像识别模型的进步对这类CAPTCHA构成重大威胁。现有方法通过添加微弱‘随机’噪声或隐藏物体来欺骗模型,但这些方法具有模型特异性,无法通用。本文发现,若允许对图像进行更显著的修改,同时保持语义信息和人类可解性,即可有效欺骗多种先进模型。具体而言,加入不同强度的遮罩后,所有模型的Accuracy @ 1(Acc@1)下降超过50个百分点,而看似鲁棒的视觉变换器(Vision Transformers)Acc@1下降达80个百分点。该结果表明,当前机器尚未真正追上人类的视觉理解能力。

原文摘要 · Abstract (English)

Modern CAPTCHAs rely heavily on vision tasks that are supposedly hard for computers but easy for humans. However, advances in image recognition models pose a significant threat to such CAPTCHAs. These models can easily be fooled by generating some well-hidden "random" noise and adding it to the image, or hiding objects in the image. However, these methods are model-specific and thus can not aid CAPTCHAs in fooling all models. We show in this work that by allowing for more significant changes to the images while preserving the semantic information and keeping it solvable by humans, we can fool many state-of-the-art models. Specifically, we demonstrate that by adding masks of various intensities the Accuracy @ 1 (Acc@1) drops by more than 50%-points for all models, and supposedly robust models such as vision transformers see an Acc@1 drop of 80%-points. These masks can therefore effectively fool modern image classifiers, thus showing that machines have not caught up with humans -- yet.

CAPTCHA对抗样本视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。