arXiv:2411.18688cs.CRcs.AI2024-11CVPR被引 52

提出推理阶段防御框架Immune,提升多模态大模型抗越狱攻击能力。

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment

  • 通过安全奖励模型控制解码,在推理时实现安全对齐。
  • 在LLaVA-1.6上使文本越狱攻击成功率降低57.82%。
  • 适合关注多模态模型安全的开发者与研究者。

随着多模态大语言模型(MLLMs)在视觉推理任务中的广泛应用,其安全性变得至关重要。研究表明,即使经过训练阶段的安全对齐,这些模型仍易受越狱攻击。本文首先指出仅靠训练阶段对齐不足以抵御越狱攻击。为此,我们提出Immune——一种推理阶段防御框架,通过安全奖励模型结合受控解码来防御越狱攻击。同时,我们提供了Immune的数学刻画,揭示其提升安全性的机制。在多个越狱基准上的评估显示,Immune有效增强模型安全性并保持原有能力。例如,在LLaVA-1.6上,相比基线模型和现有最优防御策略,其对文本越狱攻击的成功率分别降低57.82%和16.78%。

原文摘要 · Abstract (English)

With the widespread deployment of Multimodal Large Language Models (MLLMs) for visual-reasoning tasks, improving their safety has become crucial. Recent research indicates that despite training-time safety alignment, these models remain vulnerable to jailbreak attacks. In this work, we first highlight an important safety gap to describe that alignment achieved solely through safety training may be insufficient against jailbreak attacks. To address this vulnerability, we propose Immune, an inference-time defense framework that leverages a safe reward model through controlled decoding to defend against jailbreak attacks. Additionally, we provide a mathematical characterization of Immune, offering insights on why it improves safety against jailbreaks. Extensive evaluations on diverse jailbreak benchmarks using recent MLLMs reveal that Immune effectively enhances model safety while preserving the model's original capabilities. For instance, against text-based jailbreak attacks on LLaVA-1.6, Immune reduces the attack success rate by 57.82% and 16.78% compared to the base MLLM and state-of-the-art defense strategy, respectively.

多模态安全对齐越狱防御推理阶段

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。