arXiv:2502.02153cs.AIcs.CL2025-02被引 3

通过去偏推理提升安全对齐模型的帮助性,不牺牲安全性。

Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing

  • 用随机提示估计生成偏差,动态修正输出。
  • 在保持安全性的前提下,帮助性显著提升。
  • 适合关注模型公平性和可用性的研究者。

安全对齐是现实世界AI应用的关键课题。尽管现有方法普遍提升了整体安全性,但在特定类别上仍存在漏洞。我们发现,单纯减小KL惩罚、增加训练轮数或数据清洗无法有效改善安全与帮助性之间的权衡。更严重的是,安全对齐可能引发不良效应,使模型倾向于生成拒绝性回应,无论输入上下文如何。为此,我们提出无需学习的令牌级安全去偏推理(TSDI),通过随机构造提示在生成过程中估计并纠正偏差。实验表明,该方法可在维持安全性的前提下显著提升模型帮助性,优化了安全与帮助性的帕累托前沿。

原文摘要 · Abstract (English)

Safety alignment is an essential research topic for real-world AI applications. Despite the multifaceted nature of safety and trustworthiness in AI, current safety alignment methods often focus on a comprehensive notion of safety. By carefully assessing models from the existing safety-alignment methods, we found that, while they generally improved overall safety performance, they failed to ensure safety in specific categories. Our study first identified the difficulty of eliminating such vulnerabilities without sacrificing the model's helpfulness. We observed that, while smaller KL penalty parameters, increased training iterations, and dataset cleansing can enhance safety, they do not necessarily improve the trade-off between safety and helpfulness. We discovered that safety alignment could even induce undesired effects and result in a model that prefers generating negative tokens leading to rejective responses, regardless of the input context. To address this, we introduced a learning-free method, Token-level Safety-Debiased Inference (TSDI), to estimate and correct this bias during the generation process using randomly constructed prompts. Our experiments demonstrated that our method could enhance the model's helpfulness while maintaining safety, thus improving the trade-off Pareto-front.

安全对齐去偏生成优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。