arXiv:2603.22829cs.AI2026-03

解决大模型安全对齐中的过拟合问题,提升安全性能。

Improving Safety Alignment via Balanced Direct Preference Optimization

  • 基于互信息动态调节偏好响应的优化强度
  • 在多个基准上实现更强安全能力且保持通用性能
  • 适合关注大模型安全性与对齐效果的研究者

随着大语言模型(LLMs)的快速发展与广泛应用,其潜在安全风险受到广泛关注。强化学习从人类反馈(RLHF)已被用于提升LLM的安全性。作为RLHF的简单有效替代方案,直接偏好优化(DPO)被广泛用于安全对齐。然而,安全对齐仍存在严重过拟合问题,限制了实际性能。本文从模型对训练数据理解的角度重新审视过拟合现象,发现偏好对中响应之间存在不平衡的偏好理解,损害了模型安全表现。为此,提出平衡式直接偏好优化(B-DPO),根据互信息自适应调节优选与非优选响应间的优化强度。大量实验表明,相较于现有最先进方法,B-DPO在保持主流基准上竞争力的同时,显著提升了模型的安全能力。

原文摘要 · Abstract (English)

With the rapid development and widespread application of Large Language Models (LLMs), their potential safety risks have attracted widespread attention. Reinforcement Learning from Human Feedback (RLHF) has been adopted to enhance the safety performance of LLMs. As a simple and effective alternative to RLHF, Direct Preference Optimization (DPO) is widely used for safety alignment. However, safety alignment still suffers from severe overfitting, which limits its actual performance. This paper revisits the overfitting phenomenon from the perspective of the model's comprehension of the training data. We find that the Imbalanced Preference Comprehension phenomenon exists between responses in preference pairs, which compromises the model's safety performance. To address this, we propose Balanced Direct Preference Optimization (B-DPO), which adaptively modulates optimization strength between preferred and dispreferred responses based on mutual information. A series of experimental results show that B-DPO can enhance the safety capability while maintaining the competitive general capabilities of LLMs on various mainstream benchmarks compared to state-of-the-art methods. \color{red}{Warning: This paper contains examples of harmful texts, and reader discretion is recommended.

大模型安全偏好优化对齐技术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。