通过识别安全神经元,精准提升多语言视觉-语言模型的安全性。
Targeted Interpretable Safety Neuron Enhancement for Multilingual Vision-Language Large Models
- 从神经元层面识别与安全拒绝相关的激活模式。
- 仅微调少数安全神经元,显著提升跨语言跨模态安全性。
- 适合关注大模型安全对齐的开发者和研究者。
随着视觉-语言大模型(VLLMs)的广泛应用,其安全对齐面临语言与模态双重挑战。现有方法分别处理多语言与多模态安全,忽略了低资源语言指令与视觉上下文之间的耦合风险,导致难以检测跨语言、跨模态的有害意图,也影响了安全边界的稳健性。为此,我们提出一种神经元级可解释的安全对齐框架,通过对比有害请求与良性输入在前馈网络(FFN)中的表示,识别与安全拒绝对应的神经元激活强度。进一步,联合建模神经元激活与其对应下投影列,计算神经元重要性,区分通用多语言/多模态神经元与负责防御的安全神经元。最后,采用神经元定向梯度掩码,将参数更新限制在所识别神经元张成的安全子空间内,实现精确且可解释的安全增强。大量实验表明,该方法仅需微调少量安全神经元,即可显著提升多语言与多模态安全性,同时保持模型通用能力。
原文摘要 · Abstract (English)
With the widespread deployment of vision-language large models (VLLMs), their safety alignment faces dual challenges across languages and modalities. Existing methods model multilingual and multimodal safety separately, overlooking coupled risks between low-resource-language instructions and visual contexts, which hinders the detection of cross-lingual and cross-modal harmful intent and the formation of robust safety boundaries. To address this, we propose a neuron-level interpretable safety alignment framework that identifies safety neurons and performs neuron-targeted safety tuning to jointly mitigate multilingual and multimodal risks. Specifically, we compare FFN representations elicited by harmful requests and benign inputs to identify neuron activation strengths associated with safety refusals. Next, we jointly model neuron activations and corresponding down-projection columns to derive neuron-level saliency, separating general multilingual and multimodal neurons from safety neurons responsible for model defense. Finally, neuron-targeted gradient masking restricts parameter updates to the safety subspace spanned by the identified neurons, enabling precise and interpretable safety enhancement. Extensive experiments show that our method enhances multilingual and multimodal safety by tuning only a few safety neurons, while preserving general capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。