arXiv:2505.21556cs.CVcs.AI2025-05

用无毒提示诱导模型输出有害内容,暴露多模态对齐漏洞

Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts

  • 通过优化对抗图像,让模型从无毒输入生成有害回复
  • 在无毒性信号条件下仍能突破安全机制,成功率显著提升
  • 适用于黑盒场景,为安全测试提供新思路

基于优化的越狱攻击通常采用有毒延续设置,在大型视觉语言模型中遵循标准的下一个词元预测目标。在此设定下,通过优化对抗图像使模型预测出有毒提示的下一个词元。然而,我们发现该方法仅在已有毒输入时有效,当缺乏明确的毒性信号时难以诱发安全偏离。为此,我们提出一种新范式:无毒到有毒(Benign-to-Toxic, B2T)越狱。与以往工作不同,我们优化对抗图像以诱导模型从无毒条件生成有害输出。由于无毒条件本身无安全违规,图像必须单独突破模型的安全机制。所提方法在多个基准上优于现有方法,具备黑盒迁移能力,并可与文本越狱互补。这些结果揭示了多模态对齐中的一个未被充分探索的漏洞,引入了一种根本性的新越狱方向。

原文摘要 · Abstract (English)

Optimization-based jailbreaks typically adopt the Toxic-Continuation setting in large vision-language models (LVLMs), following the standard next-token prediction objective. In this setting, an adversarial image is optimized to make the model predict the next token of a toxic prompt. However, we find that the Toxic-Continuation paradigm is effective at continuing already-toxic inputs, but struggles to induce safety misalignment when explicit toxic signals are absent. We propose a new paradigm: Benign-to-Toxic (B2T) jailbreak. Unlike prior work, we optimize adversarial images to induce toxic outputs from benign conditioning. Since benign conditioning contains no safety violations, the image alone must break the model's safety mechanisms. Our method outperforms prior approaches, transfers in black-box settings, and complements text-based jailbreaks. These results reveal an underexplored vulnerability in multimodal alignment and introduce a fundamentally new direction for jailbreak approaches.

越狱攻击多模态安全对抗样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。