arXiv:2605.04261cs.CRcs.LG2026-05被引 1

攻击者用微小图像扰动让AI误判,伪造权威输出。

Laundering AI Authority with Adversarial Examples

论文配图:Laundering AI Authority with Adversarial Examples
图 1 · 摘自论文原文
  • 通过视觉对抗样本在感知层欺骗VLM,不改变模型对齐。
  • 对6个主流VLM攻击成功率22%至100%,可操纵身份与内容审核。
  • 无需新算法,旧技术即可实现,警示系统安全短板。

视觉语言模型(VLM)正被广泛用作可信权威,如社交媒体图像事实核查、产品对比和内容审核。用户默认这些系统与自己感知相同的视觉内容。本文揭示,对抗样本会破坏这一假设,实现‘AI权威洗白’:攻击者轻微扰动图像,使VLM对错误输入生成自信且权威的回应。不同于越狱或提示注入,本攻击不破坏模型对齐,仅作用于感知层面。我们证明,针对公开CLIP模型的标准攻击能可靠迁移至生产级VLM,包括GPT-5.4、Claude Opus 4.6、Gemini 3和Grok 4.2。在四个攻击场景中,该方法可放大虚假信息、诋毁个人、规避NSFW内容审查并操纵商品推荐。在数百次针对身份篡改与NSFW逃避的攻击中,六种模型的成功率介于22%至100%。无需新型攻击算法,已有十余年的基础技术即足以实现,确立了攻击者能力的下限,警示防御者关注视觉对抗鲁棒性这一现实而未解的安全问题。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are increasingly deployed as trusted authorities -- fact-checking images on social media, comparing products, and moderating content. Users implicitly trust that these systems perceive the same visual content as they do. We show that adversarial examples break this assumption, enabling \emph{AI authority laundering}: an attacker subtly perturbs an image so that the VLM produces confident and authoritative responses about the \emph{wrong} input. Unlike jailbreaks or prompt injections, our attacks do not compromise model alignment; the attack operates entirely at the perceptual level. We demonstrate that standard attacks against publicly available CLIP models transfer reliably to production VLMs -- including GPT-5.4, Claude Opus~4.6, Gemini~3, and Grok~4.2. Across four attack surfaces, we show that authority laundering can amplify misinformation, disparage individuals, evade content moderation, and manipulate product recommendations. Our attacks have high success rates: In hundreds of attacks targeting identity manipulation and NSFW evasion, we measure success rates of $22 - 100\%$ across six models. No novel attack algorithm is required: basic techniques known for over a decade suffice, establishing a lower bound on attacker capability that should concern defenders. Our results demonstrate that visual adversarial robustness is now a practical -- and still largely unsolved -- safety problem.

对抗样本AI安全视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。