arXiv:2505.17078cs.CLcs.AI2025-05被引 4

发现大模型毒性根源在全局毒化子空间,提出轻量级净化方法。

GloSS over Toxicity: Understanding and Mitigating Toxicity in LLMs via Global Toxic Subspace

  • 提出全局毒化子空间概念,比局部子空间更全面表征毒性。
  • 在多个大模型上实现顶尖去毒效果,不损失通用能力。
  • 无需大规模数据或重训练,适合实际部署场景使用。

本文研究大语言模型中毒性生成的内在机制,并提出一种高效的去毒方法。以往工作通常将前馈网络(FFN)视为毒性主要来源,将毒性区域表示为一组毒化向量或逐层子空间。然而,我们的深入分析表明,全局毒化子空间能更有效、更全面地表征模型中的毒性区域。基于此洞察,我们提出GloSS(全局毒化子空间抑制)方法,这是一种轻量级四阶段方案,通过识别并从FFN参数中移除全局毒化子空间来缓解毒性。在多种大语言模型上的实验表明,GloSS在保持模型通用能力的前提下,实现了当前最优的去毒性能,且无需大规模数据或模型重训练。

原文摘要 · Abstract (English)

This paper investigates the underlying mechanisms of toxicity generation in Large Language Models (LLMs) and proposes an effective detoxification approach. Prior work typically considers the Feed-Forward Network (FFN) as the main source of toxicity, representing toxic regions as a set of toxic vectors or layer-wise subspaces. However, our in-depth analysis reveals that the global toxic subspace offers a more effective and comprehensive representation of toxic region within the model. Building on this insight, we propose GloSS (Global Toxic Subspace Suppression), a lightweight, four-stage method that mitigates toxicity by identifying and removing the global toxic subspace from the parameters of FFN. Experiments across a range of LLMs show that GloSS achieves state-of-the-art detoxification performance while preserving the models general capabilities, without requiring large-scale data or model retraining.

大模型去毒子空间FFN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。