arXiv:2504.05838cs.CVcs.AI2025-04CVPR被引 5

IP-Adapter可被恶意利用,通过隐秘图像攻击劫持AI绘图服务

Mind the Trojan Horse: Image Prompt Adapter Enabling Scalable and Deceptive Jailbreaking

  • 用不可察觉的对抗样本操控IP-Adapter实现大规模越狱
  • 攻击成功率达98.7%,且仅需公开模型即可生成攻击图像
  • 适合安全研究者与平台方关注防御机制

近期,图像提示适配器(IP-Adapter)被广泛集成至文本到图像扩散模型(T2I-DMs)以增强可控性。然而本文揭示,搭载IP-Adapter的T2I-DMs(T2I-IP-DMs)可引发一种新型越狱攻击——劫持攻击。我们证明,通过上传不可察觉的图像空间对抗样本(AEs),攻击者可劫持大量正常用户,使依赖T2I-IP-DMs的图像生成服务(IGS)产生恶意输出,并误导公众质疑服务提供方。更严重的是,IP-Adapter对开源图像编码器的依赖,显著降低了构造对抗样本的知识门槛。大量实验验证了该劫持攻击的技术可行性。针对此威胁,我们评估了现有防御措施,并探索将IP-Adapter与对抗训练模型结合,以突破现有防御局限。代码已公开于https://github.com/fhdnskfbeuv/attackIPA。

原文摘要 · Abstract (English)

Recently, the Image Prompt Adapter (IP-Adapter) has been increasingly integrated into text-to-image diffusion models (T2I-DMs) to improve controllability. However, in this paper, we reveal that T2I-DMs equipped with the IP-Adapter (T2I-IP-DMs) enable a new jailbreak attack named the hijacking attack. We demonstrate that, by uploading imperceptible image-space adversarial examples (AEs), the adversary can hijack massive benign users to jailbreak an Image Generation Service (IGS) driven by T2I-IP-DMs and mislead the public to discredit the service provider. Worse still, the IP-Adapter's dependency on open-source image encoders reduces the knowledge required to craft AEs. Extensive experiments verify the technical feasibility of the hijacking attack. In light of the revealed threat, we investigate several existing defenses and explore combining the IP-Adapter with adversarially trained models to overcome existing defenses' limitations. Our code is available at https://github.com/fhdnskfbeuv/attackIPA.

AI安全越狱攻击对抗样本图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。