arXiv:2512.15052cs.CLcs.AI2025-12中稿 · EMNLP被引 1

通过神经元级干预,让多模态大模型自动过滤有毒输出。

SGM: Safety Glasses for Multimodal Large Language Models via Neuron-Level Detoxification

  • 识别并抑制与毒性相关的神经元,不修改模型参数。
  • 在标准和对抗性场景下将有害输出率从45.0%降至4.5%。
  • 可解释、低成本,适合需要安全生成的部署场景。

多模态大语言模型(MLLMs)虽具备多模态理解能力,但会继承弱标注预训练语料中的毒性信号,导致显性有害输出,尤其在对抗性触发下,现有无参数更新的去毒方法难以应对。本文提出SGM,一种白盒神经元级多模态干预机制,如同为有毒神经元佩戴安全眼镜:通过专家加权软抑制,重新校准一组毒性相关神经元,中和跨模态有害激活,无需任何参数更新。我们构建了MM-TOXIC-QA多模态毒性数据框架,并对比了SGM与现有去毒方法。在开源MLLM上的实验表明,SGM在标准与对抗条件下均显著降低显性毒性,平均有害率从45.0%降至4.5%,同时保持语言流畅性与多模态推理能力。SGM具有可扩展性,其组合防御方案SGM*可与现有方法融合,进一步提升安全性,提供可解释、低成本的多模态生成安全控制方案。

原文摘要 · Abstract (English)

Disclaimer: Samples in this paper may be harmful and cause discomfort. Multimodal large language models (MLLMs) enable multimodal understanding but inherit toxic signals from weakly curated pretraining corpora, leading to explicitly toxic outputs, especially under adversarial triggers that late, opaque training-free detoxification methods struggle to handle. We propose SGM, a white-box neuron-level multimodal intervention that acts like safety glasses for toxic neurons: it recalibrates a set of toxicity-associated neurons via expertise-weighted soft suppression, neutralizing harmful cross-modal activations without any parameter updates. We establish MM-TOXIC-QA, a multimodal toxicity data framework, and compare SGM with existing detoxification techniques. Experiments on open-source MLLMs show that SGM mitigates explicit toxicity in standard and adversarial conditions, cutting average harmful rates from 45.0% to 4.5% while preserving fluency and multimodal reasoning. SGM is extensible, and its combined defenses, denoted as SGM*, integrate with existing detoxification methods for stronger safety performance, providing an interpretable, low-cost solution for toxicity-controlled multimodal generation.

多模态安全控制神经元干预去毒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。