arXiv:2609.08634cs.LGcs.AI2026-09

Suan改进大模型安全对齐,解决过度拒绝与质量下降问题。

Suan: Rectifying Direct Preference Safety Alignment in Large Language Models

论文配图:Suan: Rectifying Direct Preference Safety Alignment in Large Language Models
图 1 · 摘自论文原文
  • 在梯度层面直接优化,避免传统变分推导
  • 安全对齐更优,同时保持响应实用性
  • 适合关注开源模型安全性的研究者

将稳健的安全防护机制融入大语言模型(LLMs)对于生成有益且无害的回应至关重要。尽管专有系统展现出可靠的安全部署,但其底层方法和权衡关系仍不透明。在开放权重模型中实现类似安全性仍是持续挑战,因后训练版本常出现过度拒绝和通用性能下降。为此,我们提出Suan,一种新型偏好优化算法。不同于现有方法,我们在梯度层面直接定义优化目标,跳过标准变分推导。结果获得更可解释、更鲁棒的训练动态。在多样化竞争基线与基准测试中的广泛评估表明,Suan在完全保留响应效用的前提下实现了更优的安全对齐。

原文摘要 · Abstract (English)

Integrating robust safety guardrails into Large Language Models (LLMs) is essential for delivering helpful yet harmless responses. While proprietary systems exhibit reliable safety controls, their underlying methodologies and trade-offs remain largely undisclosed. Achieving comparable security in open-weight models remains a persistent challenge, as post-trained variants frequently suffer from over-refusal and degraded general quality. To overcome these drawbacks, we introduce Suan, a novel preference optimization algorithm. Unlike existing methods, we formulate the optimization objective directly at the gradient level, bypassing the standard variational derivation. As a result, we obtain more interpretable and robust training dynamics. Extensive evaluations across a diverse suite of competitive baselines and benchmarks demonstrate that Suan achieves superior safety alignment while fully preserving response utility.

大模型安全偏好优化梯度优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。