arXiv:2412.00473cs.CV2024-12ACL被引 69

通过跨模态加密绕过视觉语言模型的伦理审查

Jailbreak Large Vision-Language Models Through Multi-Modal Linkage

  • 用文本与图像互为密钥,隐藏恶意指令
  • 在游戏制作场景下伪装攻击,成功率超97%
  • 适合研究模型安全或对抗攻击的学者

随着大型视觉语言模型(VLMs)的快速发展,其潜在滥用风险日益突出。已有研究揭示了VLMs易受越狱攻击,即精心设计的输入可诱导模型生成违反伦理法律的内容。然而,现有方法在应对GPT-4o等先进VLM时效果有限,主要因恶意内容暴露过多且缺乏隐蔽引导。本文提出新型越狱攻击框架——多模态联动(MML)攻击。受密码学启发,MML在文本与图像模态间执行加密解密过程,降低恶意信息暴露。为隐秘对齐模型输出至恶意意图,MML采用“邪恶对齐”技术,将攻击嵌入视频游戏制作场景中。全面实验表明,MML在SafeBench上成功率达97.80%,在MM-SafeBench上达98.81%,在HADES-Dataset上达99.07%。代码已开源:https://github.com/wangyu-ovo/MML。

原文摘要 · Abstract (English)

With the significant advancement of Large Vision-Language Models (VLMs), concerns about their potential misuse and abuse have grown rapidly. Previous studies have highlighted VLMs' vulnerability to jailbreak attacks, where carefully crafted inputs can lead the model to produce content that violates ethical and legal standards. However, existing methods struggle against state-of-the-art VLMs like GPT-4o, due to the over-exposure of harmful content and lack of stealthy malicious guidance. In this work, we propose a novel jailbreak attack framework: Multi-Modal Linkage (MML) Attack. Drawing inspiration from cryptography, MML utilizes an encryption-decryption process across text and image modalities to mitigate over-exposure of malicious information. To align the model's output with malicious intent covertly, MML employs a technique called "evil alignment", framing the attack within a video game production scenario. Comprehensive experiments demonstrate MML's effectiveness. Specifically, MML jailbreaks GPT-4o with attack success rates of 97.80% on SafeBench, 98.81% on MM-SafeBench and 99.07% on HADES-Dataset. Our code is available at https://github.com/wangyu-ovo/MML.

越狱攻击多模态安全漏洞模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。