用共享神经元统一多语言多模态安全对齐,仅更新0.03%参数
One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs

- 通过跨模态与跨语言交集识别共享安全神经元
- 在多个基准上超越现有方法,仅更新约0.03%参数
- 适合需要低成本部署全球安全防护的模型开发者
随着大视觉语言模型(LVLMs)在全球部署,多语言指令与视觉信息的结合使恶意攻击更加隐蔽和复杂。现有方法将语言与模态防御分离,加之安全数据稀缺与微调成本高,难以应对复合攻击。为此,我们提出基于模态与语言共享安全神经元(MLS-Neurons)的神经元级跨维度安全对齐框架。首先,通过对比有害与良性样本响应,量化单语言与单模态安全神经元的功能显著性。接着,通过各语言内单模态神经元交集,提取对视觉与文本风险均敏感的模态共享安全神经元(MS-Neurons),弥合模态间安全表征鸿沟。进一步以英语为语义锚点,交叉多语言的MS-Neurons,识别出同时具备模态与语言共享特性的安全神经元(MLS-Neurons),构成对抗复合攻击的关键防线。最终,仅更新这一极小共享神经元子集(约0.03%参数),即可将仅英语的安全监督迁移至多语言多模态场景。大量实验表明,该方法在多样化的多语言多模态安全基准上显著优于现有先进方法,同时保持通用性能。
原文摘要 · Abstract (English)
As large vision-language models (LVLMs) are deployed globally, the combination of multilingual instructions and visual information makes malicious attacks more covert and sophisticated than ever before. However, existing methods isolate language and modality defenses, which, coupled with the scarcity of safety data and high fine-tuning costs, makes it difficult for models to defend against compound attacks. To address this severe challenge, we propose a neuron-level cross-dimensional safety alignment framework driven by modality- and language-shared safety neurons (MLS-Neurons). First, we identify monolingual and unimodal safety neurons by comparing responses to harmful and benign samples, quantifying functional saliency through activation strength and downstream impact. Then, by intersecting these unimodal neurons within each language, we extract modality-shared safety neurons (MS-Neurons) responsive to both visual and textual risks, bridging the safety representation gap between modalities. Furthermore, using English as a semantic anchor, we intersect MS-Neurons across languages to identify modality- and language-shared safety neurons (MLS-Neurons), serving as key defenses against compound attacks. Finally, we update only this minimal subset of shared neurons (~0.03% of parameters), transferring English-only safety supervision to multilingual and multimodal scenarios. Extensive experiments show that our method significantly outperforms state-of-the-art approaches across diverse multilingual and multimodal safety benchmarks while preserving general utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。