arXiv:2503.10872cs.CVcs.AI2025-03被引 3

用文本锚点防御视觉语言模型的越狱攻击,单次查询即可生效。

TAIJI: Textual Anchoring for Immunizing Jailbreak Images in Vision Language Models

  • 通过关键短语锚定视觉与文本内容,识别潜在有害信息。
  • 单次查询下有效抵御越狱攻击,且不影响正常任务性能。
  • 适合需高安全性的实际场景,如自动驾驶、医疗辅助等。

视觉语言模型(VLMs)虽具备强大的推理能力,但仍易受越狱攻击影响,导致产生有害或不道德的响应。现有防御方法多为白盒方案,需访问模型参数并进行大量修改,成本高且难以在真实场景中部署。尽管已有黑盒防御方法提出,但常需施加输入限制或多次查询,限制了其在自动驾驶等安全关键任务中的应用。为此,本文提出一种新型黑盒防御框架——文本锚点免疫越狱图像(TAIJI)。TAIJI利用关键短语进行文本锚定,增强模型对视觉与文本提示中嵌入有害内容的评估与缓解能力。与现有方法不同,TAIJI在推理时仅需一次查询即可有效运作,同时保持模型在良性任务上的性能。大量实验表明,TAIJI显著提升了VLM的安全性与可靠性,为实际部署提供了高效可行的解决方案。

原文摘要 · Abstract (English)

Vision Language Models (VLMs) have demonstrated impressive inference capabilities, but remain vulnerable to jailbreak attacks that can induce harmful or unethical responses. Existing defence methods are predominantly white-box approaches that require access to model parameters and extensive modifications, making them costly and impractical for many real-world scenarios. Although some black-box defences have been proposed, they often impose input constraints or require multiple queries, limiting their effectiveness in safety-critical tasks such as autonomous driving. To address these challenges, we propose a novel black-box defence framework called \textbf{T}extual \textbf{A}nchoring for \textbf{I}mmunizing \textbf{J}ailbreak \textbf{I}mages (\textbf{TAIJI}). TAIJI leverages key phrase-based textual anchoring to enhance the model's ability to assess and mitigate the harmful content embedded within both visual and textual prompts. Unlike existing methods, TAIJI operates effectively with a single query during inference, while preserving the VLM's performance on benign tasks. Extensive experiments demonstrate that TAIJI significantly enhances the safety and reliability of VLMs, providing a practical and efficient solution for real-world deployment.

视觉语言模型安全防御越狱攻击黑盒防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。