arXiv:2502.13095cs.CVcs.CL2025-02NeurIPS被引 17

发现视觉模型会误判危险内容,提出无训练修正方法。

Understanding and Rectifying Safety Perception Distortion in VLMs

  • 识别出多模态输入导致安全判断偏差的根源
  • 新方法在不损失视觉理解能力下提升安全对齐
  • 适合关注AI安全与模型可信性的研究者

近期研究表明,引入视觉模态后,视觉-语言模型(VLMs)对有害请求和越狱攻击的敏感性显著上升,其脆弱性超过纯文本大模型。我们深入分析发现,问题根源在于多模态输入引发了一种模态诱导的激活偏移,使模型系统性高估有害输入的安全性,这种现象称为‘安全感知扭曲’。为此,我们提出无需训练的激活偏移解耦与校准方法(ShiftDC),通过分解并校准该模态诱导的激活偏移,消除模态对安全判断的影响。该方法分离并移除了与安全相关的内容,恢复了原始文本模型的内在安全对齐,同时保留了模型的视觉-语言能力。实验结果表明,ShiftDC在多个安全基准测试中显著提升对齐性能,且未损害模型实用性。

原文摘要 · Abstract (English)

Recent studies reveal that vision-language models (VLMs) become more susceptible to harmful requests and jailbreak attacks after integrating the vision modality, exhibiting greater vulnerability than their text-only LLM backbones. To uncover the root cause of this phenomenon, we conduct an in-depth analysis and identify a key issue: multimodal inputs introduce an modality-induced activation shift toward a "safer" direction compared to their text-only counterparts, leading VLMs to systematically overestimate the safety of harmful inputs. We refer to this issue as safety perception distortion. To mitigate such distortion, we propose Activation Shift Disentanglement and Calibration (ShiftDC), a training-free method that decomposes and calibrates the modality-induced activation shift to reduce the impact of modality on safety. By isolating and removing the safety-relevant component, ShiftDC restores the inherent safety alignment of the LLM backbone while preserving the vision-language capabilities of VLMs. Empirical results demonstrate that ShiftDC significantly enhances alignment performance on safety benchmarks without impairing model utility.

模型安全视觉语言模型对抗攻击对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。