arXiv:2504.01550cs.LGcs.CL2025-04ACL被引 29

通过扭曲模型表征提升大模型安全性,有效抵御越狱攻击。

Representation Bending for Large Language Model Safety

  • 在损失函数中引入激活引导机制,改变有害行为的底层表征。
  • 在多个越狱基准上将攻击成功率降低至95%以下。
  • 无需额外防御系统,适合高风险场景部署。

大型语言模型虽强大,但存在生成有害内容等安全风险,且易受对抗攻击、微调漏洞影响,尤其在高风险场景中问题突出。现有安全增强方法(如基于人类反馈的微调)针对特定威胁,泛化能力差,或需手动构建防御体系。本文提出RepBend,将推理时的激活引导思想引入基于损失的微调,从根本上扰动有害行为的模型表征,实现可扩展的安全增强。实验表明,该方法在多种越狱基准上显著优于Circuit Breaker、RMU和NPO等先进方法,攻击成功率最高降低95%,同时对模型可用性与通用能力影响极小。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have emerged as powerful tools, but their inherent safety risks - ranging from harmful content generation to broader societal harms - pose significant challenges. These risks can be amplified by the recent adversarial attacks, fine-tuning vulnerabilities, and the increasing deployment of LLMs in high-stakes environments. Existing safety-enhancing techniques, such as fine-tuning with human feedback or adversarial training, are still vulnerable as they address specific threats and often fail to generalize across unseen attacks, or require manual system-level defenses. This paper introduces RepBend, a novel approach that fundamentally disrupts the representations underlying harmful behaviors in LLMs, offering a scalable solution to enhance (potentially inherent) safety. RepBend brings the idea of activation steering - simple vector arithmetic for steering model's behavior during inference - to loss-based fine-tuning. Through extensive evaluation, RepBend achieves state-of-the-art performance, outperforming prior methods such as Circuit Breaker, RMU, and NPO, with up to 95% reduction in attack success rates across diverse jailbreak benchmarks, all with negligible reduction in model usability and general capabilities.

大模型安全越狱防御表征扭曲微调优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。