arXiv:2603.14219cs.CV2026-03Transactions of th…

通过剪枝激活隐藏的安全路径,提升视觉语言模型防越狱能力。

Safety-Potential Pruning for Enhancing Safety Prompts Against VLM Jailbreaking Without Retraining

  • 用一次性剪枝法唤醒模型中沉睡的安全参数子网
  • 在三个模型上使越狱攻击成功率降低最高22%且不影响正常性能
  • 无需重新训练,适合快速部署于现有安全提示系统

安全提示是视觉语言模型(VLMs)抵御越狱攻击的一种可解释防御层,但其效果受限于模型潜在结构的响应性。我们发现,这些提示仅持续激活少数参数,而在正常使用时大部分参数处于静默状态。这一现象支持‘安全子网假说’:VLMs 内部存在结构上独立的安全路径,但需显式刺激才会激活。为此,我们提出 Safety-Potential Pruning,一种无需重训练的一次性剪枝框架,通过移除对安全提示响应弱的权重,放大安全相关激活。在三种代表性 VLM 架构和三个越狱基准测试中,该方法相较仅使用提示将攻击成功率最高降低 22%,同时保持良好正常表现。研究揭示剪枝不仅是压缩手段,更是激发对齐相关子网的结构干预,为增强越狱防御提供了新路径。

原文摘要 · Abstract (English)

Safety prompts constitute an interpretable layer of defense against jailbreak attacks in vision-language models (VLMs); however, their efficacy is constrained by the models' latent structural responsiveness. We observe that such prompts consistently engage a sparse set of parameters that remain largely quiescent during benign use. This finding motivates the Safety Subnetwork Hypothesis: VLMs embed structurally distinct pathways capable of enforcing safety, but these pathways remain dormant without explicit stimulation. To expose and amplify these pathways, we introduce Safety-Potential Pruning, a one-shot pruning framework that amplifies safety-relevant activations by removing weights that are less responsive to safety prompts without additional retraining. Across three representative VLM architectures and three jailbreak benchmarks, our method reduces attack success rates by up to 22% relative to prompting alone, all while maintaining strong benign performance. These findings frame pruning not only as a model compression technique, but as a structural intervention to emerge alignment-relevant subnets, offering a new path to robust jailbreak resistance.

视觉语言模型安全剪枝越狱防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。