发现视觉模型在真实场景变换下的脆弱性,高效生成黑盒攻击样本。
Discovering Natural Transformation Vulnerabilities in Black-Box Vision Models

- 用语言模型搜索自然语义编辑场景,通过反馈机制优化攻击策略。
- 在ImageNet分类器上攻击成功率显著高于已有方法,且查询次数更少。
- 攻击结果可跨模型、跨图像复用,揭示可迁移的自然变换漏洞。
自然对抗样本(NAEs)表明,视觉模型在超出范数约束的现实语义变化下会失效。然而,在黑盒设置中生成NAEs仍具挑战性,因现有生成攻击依赖替代模型、学习到的攻击先验或高成本查询优化,而暴露模型漏洞的自然变换事先未知。我们提出对抗情景攻击(ASA),一种基于查询的黑盒框架,利用多模态语言模型和现代文本引导生成编辑器,在自然语言编辑场景中搜索。ASA联合探索背景、天气及材质/颜色变换,通过胜者-败者反馈机制,并使用贪婪探索器仅组合提升攻击效果的场景。在多种ImageNet分类器上,ASA的攻击成功率显著高于以往查询型生成攻击,同时所需查询次数更少,且保持良好感知质量。此外,ASA具有图像级和提示级迁移性:其生成的对抗图像对不同架构的受试模型依然有效,所发现的编辑场景亦可应用于同类别图像,甚至部分跨架构复用。这些发现表明,视觉模型对自然变换模式存在可复用的脆弱性,ASA能高效识别黑盒环境中的此类漏洞。
原文摘要 · Abstract (English)
Natural adversarial examples (NAEs) reveal that vision models can fail under realistic semantic changes beyond norm-bounded perturbations. However, generating NAEs in a black-box setting remains challenging because existing generative attacks often rely on surrogate models, learned attack priors, or costly query-based optimization, whereas the natural transformations that expose model vulnerabilities are unknown a priori. We propose \textbf{Adversarial Scenario Attack (ASA)}, a query-based black-box framework that searches over natural-language editing scenarios using a multimodal language model and a modern text-guided generative editor. ASA jointly explores background, weather, and material/color transformations through winner--loser feedback, and uses a greedy explorer to compose only attack-improving scenarios. Across diverse ImageNet classifiers, ASA achieves substantially higher attack success rates than prior query-based generative attacks while requiring fewer victim-model queries and preserving competitive perceptual quality. Moreover, ASA exhibits both image-level and prompt-level transferability: its adversarial images remain effective across victim-model architectures, while its discovered editing scenarios can be reused across same-class images and, in some cases, across architectures. These findings suggest that vision models possess reusable vulnerabilities to natural transformation patterns, which ASA can efficiently identify in a black-box setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。