发现跨模态通用安全神经元,提升多模态大模型安全性
SafeNexus: Discovering and Steering Modality-Universal Safety Neurons in MLLMs

- 通过激活模式分析定位功能特异的神经元,识别跨模态安全核心
- 抑制通用安全神经元会严重削弱安全性能,但不影响模型整体能力
- 提出两种新策略:放大激活或定向微调,显著提升防御效果
尽管大语言模型展现出良好的安全性能,但将其扩展至多模态大语言模型(MLLMs)时,其扩展的多模态能力与现有安全机制之间存在显著差距。当前防御手段多局限于特定模态,难以应对跨模态威胁。为此,我们提出SafeNexus,一种基于神经元级干预的跨模态安全对齐框架。首先,通过分析中间层激活模式并量化功能重要性,定位功能特异神经元;利用对比数据识别模态特有安全神经元(BS-Neurons),并通过靶向抑制验证其在各模态中的安全调控作用。进一步跨模态分析发现,共享于各模态的BS-Neurons构成模态通用安全神经元(US-Neurons),是抵御有害跨模态攻击的核心。实验表明,抑制这些神经元会大幅降低各模态的安全表现,但对整体效用影响甚微。基于此,我们提出两种对齐策略:激活级安全增强器与安全神经元校准器,分别通过放大激活或定向微调来提升安全性能。大量实验证明,该方法在多种模态组合的安全基准上优于现有最先进方法,同时有效保持模型实用性。
原文摘要 · Abstract (English)
Although Large Language Models (LLMs) have demonstrated promising safety performance, extending them to Multimodal Large Language Models (MLLMs) exposes a significant gap between expanded multimodal capabilities and existing safety mechanisms. Current defenses remain predominantly confined to specific modal settings, thereby limiting their robustness against broader cross-modal threats. To bridge this gap, we introduce SafeNexus, a cross-modal safety alignment framework that adopts a dedicated neuron-level intervention strategy. First, we formulate a neuron localization paradigm that identifies functionally specialized neurons by characterizing intermediate-layer activation patterns and quantifying their functional salience through importance scoring. Building upon this paradigm, we exploit contrastive data to identify modality-bound safety neurons (BS-Neurons), and validate their role in regulating safety behavior within each modality via targeted suppression. Further cross-modal analysis defines modality-universal safety neurons (US-Neurons) as the shared subset of BS-Neurons identified across individual modalities, serving as the core for defending against harmful cross-modal attacks. We observe that suppressing these neurons substantially degrades safety performance across modalities, while leaving overall utility largely unaffected. Building on these insights, we propose two safety alignment strategies: activation-level safety amplifier and safety neuron calibrator. The proposed strategies enhance model safety through two distinct routes: the former amplifies the activation magnitudes of US-Neurons, while the latter selectively calibrates them via targeted fine-tuning. Extensive experiments demonstrate that our method outperforms prevailing state-of-the-art approaches on safety benchmarks spanning diverse modality combinations, while effectively preserving utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。