发现安全对齐导致模型误拒正常请求,提出针对性缓解方法
Deactivating Refusal Triggers: Understanding and Mitigating Overrefusal in Safety Alignment
- 识别训练数据中的拒绝触发词,分析误拒成因
- 新方法在防御越狱攻击与响应正常请求间取得更好平衡
- 适合关注大模型安全对齐实用性的研究者和工程师
安全对齐旨在通过在有害请求及其拒绝回答的配对数据上进行后训练,使大语言模型拒绝有害请求。尽管该方法在工业界广泛应用,但对齐后的模型会错误拒绝正常请求的问题(过度拒绝)仍缺乏深入研究,严重影响实际可用性。本文分析了过度拒绝的产生机制:训练数据中存在引发拒绝反应的语言线索(拒绝触发词),安全对齐使模型将这些线索与拒绝响应关联。这些触发词不仅包含有害语义,也包含无害线索,导致对正常请求的误拒。基于此机制分析,我们提出一种在微调中显式考虑拒绝触发词的方法。实验证明,该方法在抵御越狱攻击与保持对正常请求响应之间取得了更优权衡,优于现有方法。
原文摘要 · Abstract (English)
Safety alignment aims to ensure that large language models (LLMs) refuse harmful requests by post-training on harmful queries paired with refusal answers. Although safety alignment is widely adopted in industry, the overrefusal problem where aligned LLMs also reject benign queries after safety alignment post-training, remains insufficiently studied. Such an issue degrades the usability of safety alignment in real-world applications. In this paper, we examine how overrefusal arises under safety alignment, and propose a mitigation strategy inspired by our findings. We define refusal triggers as linguistic cues in the training data that elicit refusal responses, safety alignment encourages LLMs to associate refusal triggers within a training sample with refusal responses, leading aligned LLMs to refuse harmful queries. However, the refusal triggers include not only harmful linguistic cues but also non-harmful cues, therefore causing overrefusal to benign queries. Building on this mechanistic analysis, we propose a method that explicitly considers refusal triggers in the safety alignment fine-tuning. Empirical results demonstrate that our approach achieves a more favorable trade-off between defense against jailbreak attacks and responsiveness to benign queries, outperforming prior methods. Warning: this paper contains harmful and biased sentences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。