发现大模型生成有害内容有特定参数机制,可精准清除而不影响正常功能。
Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types
- 通过剪枝定位支持有害响应的关键稀疏参数
- 剪除这些参数后有害响应下降90%以上,正常能力仅轻微下降
- 该机制在不同有害类型间共享,适合安全增强研究者参考
大型语言模型仍易受越狱攻击,引发有害输出,但其内在机制尚不明确。本文从参数层面解析有害响应的生成机制,识别并剪枝了专门支持有害合规的参数。实验发现,有害响应依赖于一组稀疏的关键参数:剪除后有害合规率显著下降,而良性能力仅轻微退化,表明有害生成组件与通用能力可分离。来自某一有害类型的关键参数亦能抑制其他类型的有害输出,说明跨危害类型的共享机制存在。该可分离性主要出现在对齐训练后的模型中,表明对齐训练虽未完全消除风险,但已重构有害响应机制。此外,有害生成能力与识别危害性的能力可分离。研究还扩展至新兴的非对齐现象,发现其同样由少量稀疏参数驱动,且在不同微调领域间高度共享。结果揭示了不安全行为背后的统一参数组织,为更系统的模型安全干预提供依据。
原文摘要 · Abstract (English)
Large language models remain vulnerable to jailbreaks that elicit harmful responses, yet the mechanism behind harmful response generation is poorly understood. Here, we investigate how this capability is organized within model parameters. We identify and prune parameters that specifically support harmful compliance, providing a direct mechanistic analysis at the parameter level. We find that this capability depends on a sparse set of critical parameters: pruning these parameters substantially reduces harmful compliance while causing only limited degradation in benign capabilities, suggesting that key components of harmful generation are separable from those of general utility. Parameters identified from one harm category also reduce harmful responses in others, indicating components shared across harm types. This separability appears primarily in aligned models, suggesting that alignment training internally reshapes the harmful response mechanism even when behavioral safeguards remain brittle. We further show that harmful response generation is dissociable from the ability to recognize and reason about harmfulness. Finally, we extend our analysis to emergent misalignment and identify a sparse set of parameters contributing to it, with substantial sharing across fine-tuning domains. Together, these results reveal a consistent parameter-level organization underlying unsafe behaviors and point toward more principled interventions for improving model safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。