针对制爆指令,新方法比现有防御更有效但仍有漏洞。
Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
- 用对话记录分类器识别潜在制爆意图
- 在制爆场景下准确率超传统方法15%以上
- 适合安全团队测试大模型风险边界
防止大语言模型生成广泛禁止内容仍是开放问题。本文聚焦于狭窄行为限制:阻止模型协助用户制造炸弹。研究发现,主流防御方法如安全训练、对抗训练和输入/输出分类器均无法完全解决该问题。为此,我们提出一种基于对话记录的分类器防御机制,在测试中优于现有基线方案。然而,该方法仍存在失效情况,表明即使在狭义领域内,防范越狱攻击依然极具挑战性。
原文摘要 · Abstract (English)
Defending large language models against jailbreaks so that they never engage in a broadly-defined set of forbidden behaviors is an open problem. In this paper, we investigate the difficulty of jailbreak-defense when we only want to forbid a narrowly-defined set of behaviors. As a case study, we focus on preventing an LLM from helping a user make a bomb. We find that popular defenses such as safety training, adversarial training, and input/output classifiers are unable to fully solve this problem. In pursuit of a better solution, we develop a transcript-classifier defense which outperforms the baseline defenses we test. However, our classifier defense still fails in some circumstances, which highlights the difficulty of jailbreak-defense even in a narrow domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。