arXiv:2502.14486cs.CRcs.AI2025-02

剖析越狱防御机制,提出双策略平衡安全与可用性。

How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation

  • 将生成任务转为二分类,量化模型拒绝有害请求的能力。
  • 发现两种核心防御机制:安全偏移与有害性区分,可提升拒绝率。
  • 设计跨机制与同机制集成策略,显著优化安全与有用性权衡。

越狱攻击使有害提示绕过生成模型的安全机制,引发严重隐患。尽管已有多种防御方法,但其在安全性与实用性间的权衡,以及在大视觉语言模型(LVLMs)中的应用仍不清晰。本文通过将标准生成任务重构为二分类问题,评估模型对有害与良性查询的拒绝倾向。识别出两种关键防御机制:安全偏移(提升所有查询的拒绝率)与有害性区分(增强模型辨别能力)。基于此,提出两类集成防御策略——跨机制集成与同机制集成,以平衡安全与帮助性。在MM-SafetyBench和MOSSBench数据集上,使用LLaVA-1.5模型的实验表明,这些策略能有效提升模型安全性或优化安全与可用性之间的权衡。

原文摘要 · Abstract (English)

Jailbreak attacks, where harmful prompts bypass generative models' built-in safety, raise serious concerns about model vulnerability. While many defense methods have been proposed, the trade-offs between safety and helpfulness, and their application to Large Vision-Language Models (LVLMs), are not well understood. This paper systematically examines jailbreak defenses by reframing the standard generation task as a binary classification problem to assess model refusal tendencies for both harmful and benign queries. We identify two key defense mechanisms: safety shift, which increases refusal rates across all queries, and harmfulness discrimination, which improves the model's ability to distinguish between harmful and benign inputs. Using these mechanisms, we develop two ensemble defense strategies-inter-mechanism ensembles and intra-mechanism ensembles-to balance safety and helpfulness. Experiments on the MM-SafetyBench and MOSSBench datasets with LLaVA-1.5 models show that these strategies effectively improve model safety or optimize the trade-off between safety and helpfulness.

越狱防御安全评估集成策略视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。