通过干预模型隐空间,有效抑制有毒内容生成,不牺牲文本流畅性。
Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention
- 在模型隐空间中定位并压制隐藏的毒性模式
- 对主流去毒系统提升8%-20%毒性缓解效果
- 适合关注大模型安全、需兼顾生成质量的研究者
大型语言模型(LLMs)虽具强大文本生成能力,却可能在看似无害的提示下产生有害内容,带来现实风险。毒性常具隐蔽性和语境依赖性,难以通过词元或粗粒度句级信号检测。现有去毒方法常面临安全性与文本连贯性之间的权衡。本文提出一种针对性的隐空间干预策略,旨在识别并抑制模型表征中的隐藏毒性模式,同时保持生成安全流畅内容的能力。在RealToxicityPrompts数据集上,该方法相较现有基线表现优异,且推理复杂度影响极小。在多个LLM上,其使当前最优去毒系统毒性降低8%-20%,同时保持相近流畅性。通过大量定量与定性分析表明,该方法能有效降低毒性而不损害生成性能,持续优于现有基线。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are powerful text generators, yet they can produce toxic or harmful content even when given seemingly harmless prompts. This presents a serious safety challenge and can cause real-world harm. Toxicity is often subtle and context-dependent, making it difficult to detect at the token level or through coarse sentence-level signals. Moreover, efforts to mitigate toxicity often face a trade-off between safety and the coherence, or fluency of the generated text. In this work, we present a targeted subspace intervention strategy for identifying and suppressing hidden toxic patterns from underlying model representations, while preserving overall ability to generate safe fluent content. On the RealToxicityPrompts, our method achieves strong mitigation performance compared to existing baselines, with minimal impact on inference complexity. Across multiple LLMs, our approach reduces toxicity of state-of-the-art detoxification systems by 8-20%, while maintaining comparable fluency. Through extensive quantitative and qualitative analyses, we show that our approach achieves effective toxicity reduction without impairing generative performance, consistently outperforming existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。