arXiv:2411.18000cs.CV2024-11CVPR被引 12

提出多损失对抗搜索框架,突破视觉语言模型安全防护瓶颈。

Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models

  • 基于场景感知生成图像,利用平坦极小值理论选择抗干扰攻击图
  • 在MiniGPT-4上实现77.75%攻击成功率,比现有方法高34.37%
  • 适用于黑盒商用模型,可迁移至真实场景防御研究

尽管继承了底层语言模型的安全机制,视觉语言模型(VLM)仍可能面临安全对齐问题。通过实证分析,我们发现:场景匹配的图像会显著放大有害输出;且与梯度攻击的常见假设相反,损失值越小并不意味着攻击效果越好。基于这些发现,我们提出MLAI(多损失对抗图像)框架,利用场景感知图像生成实现语义对齐,借助平坦极小值理论选择鲁棒对抗图像,并采用多图像协同攻击提升有效性。大量实验表明,MLAI在MiniGPT-4上达到77.75%的攻击成功率,在LLaVA-2上达82.80%,分别优于现有方法34.37%和12.77%。此外,该方法在商业黑盒VLM中表现出显著迁移性,最高成功率可达60.11%。本工作揭示了当前VLM安全机制中的根本视觉漏洞,强调需加强防御措施。警告:本文包含潜在有害示例文本。

原文摘要 · Abstract (English)

Despite inheriting security measures from underlying language models, Vision-Language Models (VLMs) may still be vulnerable to safety alignment issues. Through empirical analysis, we uncover two critical findings: scenario-matched images can significantly amplify harmful outputs, and contrary to common assumptions in gradient-based attacks, minimal loss values do not guarantee optimal attack effectiveness. Building on these insights, we introduce MLAI (Multi-Loss Adversarial Images), a novel jailbreak framework that leverages scenario-aware image generation for semantic alignment, exploits flat minima theory for robust adversarial image selection, and employs multi-image collaborative attacks for enhanced effectiveness. Extensive experiments demonstrate MLAI's significant impact, achieving attack success rates of 77.75% on MiniGPT-4 and 82.80% on LLaVA-2, substantially outperforming existing methods by margins of 34.37% and 12.77% respectively. Furthermore, MLAI shows considerable transferability to commercial black-box VLMs, achieving up to 60.11% success rate. Our work reveals fundamental visual vulnerabilities in current VLMs safety mechanisms and underscores the need for stronger defenses. Warning: This paper contains potentially harmful example text.

视觉安全对抗攻击模型漏洞

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。