用模型自身语义理解提升视觉语言模型安全性,无需微调
Self-Aware Safety Augmentation: Leveraging Internal Semantic Understanding to Enhance Safety in Vision-Language Models
- 从中间层提取语义信息,注入早期安全层增强识别能力
- 在多个数据集上显著提升安全性,对模型性能影响极小
- 适合关注大模型安全性的研究者与开发者
大型视觉语言模型(LVLMs)相较于纯语言模型更易受有害输入影响。我们通过分析模型内部动态,将内在安全理解归纳为三个核心能力:安全感知、语义理解与语言表达对齐,并实验定位了这些能力在模型架构中的主要分布位置。结果表明,安全感知常在全面语义理解之前出现,导致安全能力下降。为此,我们提出自洽式安全增强(SASA)技术,将中间层的语义表示投影到早期安全导向层,利用模型自身的语义理解能力强化安全识别,无需微调。同时采用线性探测揭示模型内部语义认知,实现生成前的风险检测。在多种数据集和任务上的大量实验表明,SASA显著提升了LVLMs的安全性,且对模型实用性影响极小。
原文摘要 · Abstract (English)
Large vision-language models (LVLMs) are vulnerable to harmful input compared to their language-only backbones. We investigated this vulnerability by exploring LVLMs internal dynamics, framing their inherent safety understanding in terms of three key capabilities. Specifically, we define these capabilities as safety perception, semantic understanding, and alignment for linguistic expression, and experimentally pinpointed their primary locations within the model architecture. The results indicate that safety perception often emerges before comprehensive semantic understanding, leading to the reduction in safety. Motivated by these findings, we propose \textbf{Self-Aware Safety Augmentation (SASA)}, a technique that projects informative semantic representations from intermediate layers onto earlier safety-oriented layers. This approach leverages the model's inherent semantic understanding to enhance safety recognition without fine-tuning. Then, we employ linear probing to articulate the model's internal semantic comprehension to detect the risk before the generation process. Extensive experiments on various datasets and tasks demonstrate that SASA significantly improves the safety of LVLMs, with minimal impact on the utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。