arXiv:2603.03323cs.CLcs.AI2026-03被引 1

通过对比精炼提升模型辨别真伪有害提示的能力

Discern Truth from Falsehood: Reducing Over-Refusal via Contrastive Refinement

  • 引入对比精炼阶段,强化模型区分真实与表面有害内容的能力
  • 在多个基准上减少过拒率,同时保持安全防护效果
  • 适合关注安全对齐中过度拒绝问题的研究者和开发者

对齐安全的大型语言模型常出现过拒现象,即误将看似有害或无害的提示判定为有害,削弱了模型的实用性。现有缓解方法如数据增强和激活引导常面临权衡:降低过拒率会损害模型对真正有害内容的识别能力。本文认为该问题源于有害与表面有害提示对学习动态的模糊影响。为此提出前置对齐阶段DCR(Discernment via Contrastive Refinement),理论上与实证上均表明对比精炼能提升模型辨别真实有害与表面有害提示的能力。跨多个基准评估显示,该方法有效降低过拒率,同时保留对齐带来的安全收益,且对通用能力影响极小,为安全对齐提供了更稳健、更根本的方向。

原文摘要 · Abstract (English)

Large language models (LLMs) aligned for safety often suffer from over-refusal, the tendency to reject seemingly toxic or benign prompts by misclassifying them as toxic. This behavior undermines models' helpfulness and restricts usability in sensitive or nuanced contexts. While prior work has proposed mitigation strategies such as data augmentation and activation steering, these approaches often face a trade-off: reducing over-refusal typically degrades the model's ability to reject genuinely harmful content. We argue that this issue arises from the ambiguous influence of toxic and seemingly toxic prompts on the model's learning dynamics. To address it, we introduce a preceding alignment stage, DCR: Discernment via Contrastive Refinement. Both theoretically and empirically, we demonstrate that contrastive refinement improves an LLM's capacity to distinguish truly toxic prompts from superficially toxic ones. Evaluation across diverse benchmarks shows that our method effectively reduces over-refusal while preserving the safety benefits of alignment. Importantly, it achieves this with minimal degradation of general capabilities, offering a more principled and robust direction for safety alignment.

大模型安全过拒问题对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。