用偏好优化让大模型更安全,效果接近顶尖水平。
Alignment with Preference Optimization Is All You Need for LLM Safety
- 通过偏好优化提升模型安全性,使用安全数据集微调。
- 安全评分从57.64%升至99.90%,毒性检测得分低于0.07。
- 推荐使用Safe-NCA方法,在安全与能力间取得平衡。
我们证明了偏好优化方法能有效提升大模型的安全性。在Falcon 11B模型上应用多种对齐技术,并结合安全数据集进行微调,其在LlamaGuard 3 8B评估下的全局安全分数从57.64%显著提升至99.90%,达到当前领先水平。在对抗性毒性和有害内容检测任务中,平均得分从超过0.6降至不足0.07。然而,这一安全提升以牺牲部分通用能力为代价,尤其在数学推理方面表现下降,表明存在安全与性能的权衡。研究发现噪声对比对齐(Safe-NCA)是平衡安全与性能的最佳方法。结果表明,仅通过合适的对齐技术即可构建出安全且鲁棒的大模型。
原文摘要 · Abstract (English)
We demonstrate that preference optimization methods can effectively enhance LLM safety. Applying various alignment techniques to the Falcon 11B model using safety datasets, we achieve a significant boost in global safety score (from $57.64\%$ to $99.90\%$) as measured by LlamaGuard 3 8B, competing with state-of-the-art models. On toxicity benchmarks, average scores in adversarial settings dropped from over $0.6$ to less than $0.07$. However, this safety improvement comes at the cost of reduced general capabilities, particularly in math, suggesting a trade-off. We identify noise contrastive alignment (Safe-NCA) as an optimal method for balancing safety and performance. Our study ultimately shows that alignment techniques can be sufficient for building safe and robust models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。