arXiv:2501.11183cs.CRcs.AI2025-01被引 1

安全微调像网络安全攻防战,需从设计源头构建更可靠的防护机制。

Can Safety Fine-Tuning Be More Principled? Lessons Learned from Cybersecurity

  • 将安全微调类比为攻防对抗,指出现有方法治标不治本
  • 实证显示当前防御易被新型越狱攻击、奖励滥用等绕过
  • 提出应借鉴安全工程经验,从架构上构建可信赖的模型

随着大语言模型能力不断增强,其可能对社会造成的危害也日益凸显,因此多数模型通过微调添加安全防护。本文认为,当前的安全微调本质上类似于网络安全中的攻防博弈(或军备竞赛):针对具体攻击手法进行修补,但诸多相似攻击路径仍存。当防御方缺乏系统性原则时,攻击者极易规避新防御。我们证明现有防护无法有效阻止新型对抗性越狱攻击、奖励劫持及失控问题。为吸取网络安全历史教训,本文引入经典案例类比,提炼出可应用于大模型安全的若干经验。这些观点支持建立更系统、更根本性的安全设计范式,即从模型架构之初就融入安全性考量。文中还介绍了人工智能领域中若干具备此特性的方法。

原文摘要 · Abstract (English)

As LLMs develop increasingly advanced capabilities, there is an increased need to minimize the harm that could be caused to society by certain model outputs; hence, most LLMs have safety guardrails added, for example via fine-tuning. In this paper, we argue the position that current safety fine-tuning is very similar to a traditional cat-and-mouse game (or arms race) between attackers and defenders in cybersecurity. Model jailbreaks and attacks are patched with bandaids to target the specific attack mechanism, but many similar attack vectors might remain. When defenders are not proactively coming up with principled mechanisms, it becomes very easy for attackers to sidestep any new defenses. We show how current defenses are insufficient to prevent new adversarial jailbreak attacks, reward hacking, and loss of control problems. In order to learn from past mistakes in cybersecurity, we draw analogies with historical examples and develop lessons learned that can be applied to LLM safety. These arguments support the need for new and more principled approaches to designing safe models, which are architected for security from the beginning. We describe several such approaches from the AI literature.

安全微调越狱攻击模型安全攻防对抗

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。