arXiv:2412.08156cs.CRcs.AI2024-12被引 3

用语义混淆绕过安全检测,隐蔽生成违规图像

Antelope: Potent and Concealed Jailbreak Attack Strategy

  • 通过混淆敏感概念与相似概念,搜索语义邻近空间
  • 在多类防御机制下仍能高效生成违规内容
  • 可迁移攻击线上黑盒服务,适合安全漏洞研究

由于基于扩散模型的强大生成能力,针对此类框架的越狱攻击受到广泛关注。图像模型中一个特别令人担忧的问题是生成不适宜工作场所(NSFW)内容。尽管已部署安全过滤机制,仍有不少研究尝试绕过这些防护。现有攻击方法主要依赖对抗性提示工程或概念模糊化,但常存在搜索效率低、攻击特征明显、与目标对齐差等问题。为此,我们提出Antelope,一种更鲁棒且隐蔽的越狱攻击策略,旨在揭示生成模型内在的安全漏洞。具体而言,Antelope利用敏感概念与相似概念之间的混淆,在相关概念的语义邻近空间中进行搜索,并使其与目标图像对齐,从而生成与目标一致且可逃避检测的敏感图像。此外,我们成功利用模型攻击的可迁移性,突破了在线黑盒服务的限制。实验评估表明,Antelope在多种防御机制下均优于现有基线,凸显其有效性和通用性。

原文摘要 · Abstract (English)

Due to the remarkable generative potential of diffusion-based models, numerous researches have investigated jailbreak attacks targeting these frameworks. A particularly concerning threat within image models is the generation of Not-Safe-for-Work (NSFW) content. Despite the implementation of security filters, numerous efforts continue to explore ways to circumvent these safeguards. Current attack methodologies primarily encompass adversarial prompt engineering or concept obfuscation, yet they frequently suffer from slow search efficiency, conspicuous attack characteristics and poor alignment with targets. To overcome these challenges, we propose Antelope, a more robust and covert jailbreak attack strategy designed to expose security vulnerabilities inherent in generative models. Specifically, Antelope leverages the confusion of sensitive concepts with similar ones, facilitates searches in the semantically adjacent space of these related concepts and aligns them with the target imagery, thereby generating sensitive images that are consistent with the target and capable of evading detection. Besides, we successfully exploit the transferability of model-based attacks to penetrate online black-box services. Experimental evaluations demonstrate that Antelope outperforms existing baselines across multiple defensive mechanisms, underscoring its efficacy and versatility.

越狱攻击扩散模型安全漏洞语义混淆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。