arXiv:2410.01438cs.LG2024-10被引 1

揭示视觉语言模型中攻击效果与隐蔽性的根本矛盾

The Great Contradiction Showdown: How Jailbreak and Stealth Wrestle in Vision-Language Models?

  • 基于信息论建立攻击效果与隐蔽性的理论框架
  • 实验证明强攻击易被检测,存在可防御的平衡点
  • 提出高效检测非隐蔽攻击的算法,提升模型鲁棒性

视觉语言模型(VLMs)在多项任务中表现优异,但仍易受越狱攻击影响,威胁安全与可靠性。本文提出一种信息论框架,揭示攻击有效性与隐蔽性之间的内在权衡。基于法诺不等式,证明攻击成功率与提示隐蔽性紧密相关。在此基础上,设计了一种高效算法,用于检测非隐蔽的越狱攻击,显著提升模型鲁棒性。实验结果凸显强攻击与可检测性之间的张力,为对抗策略与防御机制提供了深入洞见。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have achieved remarkable performance on a variety of tasks, yet they remain vulnerable to jailbreak attacks that compromise safety and reliability. In this paper, we provide an information-theoretic framework for understanding the fundamental trade-off between the effectiveness of these attacks and their stealthiness. Drawing on Fano's inequality, we demonstrate how an attacker's success probability is intrinsically linked to the stealthiness of generated prompts. Building on this, we propose an efficient algorithm for detecting non-stealthy jailbreak attacks, offering significant improvements in model robustness. Experimental results highlight the tension between strong attacks and their detectability, providing insights into both adversarial strategies and defense mechanisms.

视觉语言模型越狱攻击信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。