arXiv:2502.00657cs.LGcs.AI2025-02NeurIPS被引 7

将大模型对齐视为分布差异估计,揭示安全与有害提示的潜在空间分离机制。

LLM Safety Alignment is Divergence Estimation in Disguise

  • 把对齐方法看作安全与有害分布间的差异估计器
  • 提出基于KL散度的新方法KLDO,实证有效提升对齐效果
  • 用合规拒绝数据集可增强分离度,适合安全敏感场景

我们提出一个理论框架,表明主流的大语言模型对齐方法(如RLHF及其变体)本质上是安全(理想)与不安全(有害)分布间差异的估计器。该视角解释了对齐后潜在空间中安全与有害提示出现分离现象的原因。作为该通用框架的应用,我们提出了基于KL散度的新型对齐方法KLDO,并通过实验验证其有效性。进一步研究表明,使用合规-拒绝数据集而非标准偏好数据集,能带来更强的分离效果并提升安全性。最后,我们提出一种基于距离的提示表示空间度量,该指标在统计上显著关联模型安全性。

原文摘要 · Abstract (English)

We present a theoretical framework showing that popular LLM alignment methods, including RLHF and its variants, can be understood as divergence estimators between aligned (safe or preferred) and unaligned (harmful or less preferred) distributions. This perspective explains the emergence of separation in the latent space between safe and harmful prompts after alignment. As an application of our general divergence framework, we propose KLDO, a novel KL divergence-based alignment method, and empirically validate its effectiveness. We further show that using compliance-refusal datasets, rather than standard preference-based datasets, leads to stronger separation and improved safety alignment. Finally, to quantify the separation effect, we propose a distance-based metric in the prompt representation space, which also acts as a statistically significant indicator for model safety.

大模型对齐分布差异安全性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。