arXiv:2511.19009cs.CRcs.CL2025-11

通过干预表示空间缓解大模型过度拒绝问题。

Understanding and Mitigating Over-refusal for Large Language Models via Representation Intervention

  • 从表示层面分析过度拒绝成因,发现模型难区分恶意与误拒样本。
  • 提出重叠感知损失加权与上下文感知增强策略,降低误拒率。
  • 兼顾安全防御与可用性,适合关注模型可靠性研究者参考。

大语言模型在自然语言处理任务中表现强大,但其内在安全漏洞制约了实际应用的可靠性。尽管已有多种越狱防御方法提升安全性,却常伴随严重过度拒绝问题,难以平衡安全与可用性。本文从表示视角分析过度拒绝成因,发现模型无法有效区分过度拒绝样本与恶意样本。为此,提出在安全表示空间进行干预的方法:(1) 重叠感知损失加权,通过量化恶意样本与过度拒绝样本在表示空间中的相似性来确定擦除权重;(2) 上下文感知增强,在拒绝响应前添加有害前缀以补充决策上下文。实验表明,该方法在缓解过度拒绝的同时保持良好安全性,优于现有方法。本文也呼吁研究者从安全与过度拒绝双重角度评估防御方法的可靠性。

原文摘要 · Abstract (English)

Large language models (LLMs) demonstrate powerful capabilities across various natural language processing tasks,yet their inherent safety vulnerabilities undermine the reliable application of LLMs in real-world scenarios. To enhance LLM safety, various jailbreak defense methods have been proposed to guard against harmful outputs. However, improvements in model safety often come at the cost of severe over-refusal, failing to strike a good balance between safety and usability. This phenomenon is a critical reliability degradation issue in LLM intelligent systems, failing to strike a good balance between safety defense effectiveness and system usability reliability. In this paper, we first analyze the causes of over-refusal from a representation perspective, revealing that LLMs are unable to effectively distinguish between over-refusal samples and malicious samples. Based on this, we propose to mitigate overrefusal by intervening in the safety representation space of LLMs. Our method incorporates two core strategies: (1) OverlapAware Loss Weighting, which determines the erasure weight for malicious samples by quantifying their similarity to overrefusal samples in the representation space, and (2) ContextAware Augmentation, which supplements the necessary context for rejection decisions by adding harmful prefixes before rejection responses. Experiments demonstrate that our method achieves a better trade-off between mitigating over-refusal and maintaining safety, compared with existing approaches. This paper also aims to encourage researchers to consider the reliability of defending methods against jailbreak attacks from both the perspectives of safety and over-refusal.

大模型安全过度拒绝表示干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。