arXiv:2409.20089cs.LGcs.CL2024-09ICLR被引 73

通过模拟拒绝特征扰动,高效提升大模型抗攻击能力。

Robust LLM safeguarding via refusal feature adversarial training

  • 通过擦除残差流中的拒绝特征来模拟攻击
  • 显著增强三款主流大模型的防御效果
  • 计算开销远低于传统对抗训练方法

大型语言模型易受对抗攻击,诱导产生有害响应。由于越狱机制不透明且训练鲁棒模型计算成本高,防御仍具挑战。我们发现对抗攻击共享一种通用机制:通过擦除残差流嵌入空间中的一个维度——拒绝特征(refusal feature),即可绕过模型安全防护。进一步表明,拒绝特征擦除(RFA)操作近似于使模型安全性的最坏情况扰动。基于此,我们提出拒绝特征对抗训练(ReFAT),一种通过模拟输入级攻击效应实现高效对抗训练的新算法。实验显示,ReFAT显著提升了三款主流大模型对多种对抗攻击的鲁棒性,且计算开销远低于现有方法。

原文摘要 · Abstract (English)

Large language models (LLMs) are vulnerable to adversarial attacks that can elicit harmful responses. Defending against such attacks remains challenging due to the opacity of jailbreaking mechanisms and the high computational cost of training LLMs robustly. We demonstrate that adversarial attacks share a universal mechanism for circumventing LLM safeguards that works by ablating a dimension in the residual stream embedding space called the refusal feature. We further show that the operation of refusal feature ablation (RFA) approximates the worst-case perturbation of offsetting model safety. Based on these findings, we propose Refusal Feature Adversarial Training (ReFAT), a novel algorithm that efficiently performs LLM adversarial training by simulating the effect of input-level attacks via RFA. Experiment results show that ReFAT significantly improves the robustness of three popular LLMs against a wide range of adversarial attacks, with considerably less computational overhead compared to existing adversarial training methods.

大模型安全对抗训练拒绝特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。