提出全局毒性子空间抑制法,高效净化大模型毒性和风险内容。
Projecting Out the Malice: A Global Subspace Approach to LLM Detoxification
- 通过识别并移除全连接层中的全局毒性子空间来净化模型
- 在不重训练情况下实现最优去毒效果,保持模型通用能力
- 适合关注大模型安全、内容过滤与可信生成的研究者
大型语言模型虽性能优异,但易生成有害内容,限制其安全应用。传统对齐方法仅调整输出偏好,无法消除参数中深层的毒性区域,仍易受对抗攻击。已有研究将毒性区域视为“毒性向量”或“逐层子空间”,但分析发现:1)移除的毒性向量可由非毒性向量线性重构,必须针对整个毒性子空间;2)基于有限样本的对比目标会引入噪声,干扰子空间稳定提取。为此,本文提出轻量级方法 GLOSS(GLobal tOxic Subspace Suppression),通过识别并清除前馈网络参数中的全局毒性子空间实现去毒。在 Qwen3 等模型上的实验表明,该方法在无需大规模重训练的前提下达到当前最优去毒效果,同时保留模型通用能力。
原文摘要 · Abstract (English)
Large language models (LLMs) exhibit exceptional performance but pose inherent risks of generating toxic content, restricting their safe deployment. While traditional methods (e.g., alignment) adjust output preferences, they fail to eliminate underlying toxic regions in parameters, leaving models vulnerable to adversarial attacks. Prior mechanistic studies characterize toxic regions as "toxic vectors" or "layer-wise subspaces", yet our analysis identifies critical limitations: i) Removed toxic vectors can be reconstructed via linear combinations of non-toxic vectors, demanding targeting of entire toxic subspace; ii) Contrastive objective over limited samples inject noise into layer-wise subspaces, hindering stable extraction. These highlight the challenge of identifying robust toxic subspace and removing them. Therefore, we propose GLOSS (GLobal tOxic Subspace Suppression), a lightweight method that mitigates toxicity by identifying and eliminating this global subspace from FFN parameters. Experiments on LLMs (e.g., Qwen3) show GLOSS achieves SOTA detoxification while preserving general capabilities without requiring large-scale retraining. WARNING: This paper contains context which is toxic in nature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。