arXiv:2603.22094cs.CV2026-03被引 4

通过投影到零空间,实现模型安全与性能的平衡防御。

Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language Models

  • 在激活空间中构建拒绝方向,仅对有害输入响应。
  • 平均降低15%以上攻击成功率,保持原有任务表现。
  • 理论可解释,适合部署于高风险视觉语言场景。

随着视觉语言模型(VLMs)在开放世界场景中的广泛应用,其易受视觉越狱攻击的影响,导致生成有害内容,威胁模型安全与可信使用。现有激活控制方法在推理时注入方向向量以诱导拒绝行为,虽有效但常引发过度拒绝,损害正常输入表现。此外,由于缺乏理论可解释性,这些方法鲁棒性与有效性受限。为此,我们提出NullSteer——一种基于零空间投影的激活防御框架。该方法通过线性变换在模型激活中构建拒绝方向:在良性子空间内保持零扰动,同时动态诱导潜在有害方向上的拒绝,理论上实现安全增强而不影响模型通用能力。大量实验表明,NullSteer在多种越狱攻击下显著减少有害输出(在MiniGPT-4上平均攻击成功率降低超15%),且在通用基准上性能接近原始模型。

原文摘要 · Abstract (English)

As vision-language models (VLMs) are increasingly deployed in open-world scenarios, they can be easily induced by visual jailbreak attacks to generate harmful content, posing serious risks to model safety and trustworthy usage. Recent activation steering methods inject directional vectors into model activations during inference to induce refusal behaviors and have demonstrated effectiveness. However, a steering vector may both enhance refusal ability and cause over-refusal, thereby degrading model performance on benign inputs. Moreover, due to the lack of theoretical interpretability, these methods still suffer from limited robustness and effectiveness. To better balance safety and utility, we propose NullSteer, a null-space projected activation defense framework. Our method constructs refusal directions within model activations through a linear transformation: it maintains zero perturbation within the benign subspace while dynamically inducing refusal along potentially harmful directions, thereby theoretically achieving safety enhancement without impairing the model's general capabilities. Extensive experiments show that NullSteer significantly reduces harmful outputs under various jailbreak attacks (average ASR reduction over 15 percent on MiniGPT-4) while maintaining comparable performance to the original model on general benchmarks.

模型安全越狱防御视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。