利用无限改写攻击突破大模型安全防线,威胁商业级AI应用。
Jailbreaking Large Language Models in Infinitely Many Ways
- 通过语义映射和编码实现无限改写,绕过模型防御机制。
- 可成功攻破最强大开源与闭源大模型的安全策略。
- 适合关注大模型安全漏洞的研究者与产品开发者。
我们探讨了"无限改写攻击"(IMP),这类攻击利用模型对改写和编码通信能力的提升,绕过其安全防护机制。IMP的有效性随模型能力增强而增长,在实践中表现极佳,对商业级最强大大语言模型构成切实威胁。本文展示如何突破最先进开放与闭源模型的安全限制,生成明确违反安全政策的内容。可通过增强并随模型能力同步升级的护栏来抵御IMP。针对两种易实现的攻击类型——双射与编码,我们提出分别在词元和嵌入空间中的防御策略。最后,我们提出若干应优先研究的问题,以提升大模型的防御能力并深化对其安全性的理解。
原文摘要 · Abstract (English)
We discuss the ``Infinitely Many Paraphrases'' attacks (IMP), a category of jailbreaks that leverages the increasing capabilities of a model to handle paraphrases and encoded communications to bypass their defensive mechanisms. IMPs' viability pairs and grows with a model's capabilities to handle and bind the semantics of simple mappings between tokens and work extremely well in practice, posing a concrete threat to the users of the most powerful LLMs in commerce. We show how one can bypass the safeguards of the most powerful open- and closed-source LLMs and generate content that explicitly violates their safety policies. One can protect against IMPs by improving the guardrails and making them scale with the LLMs' capabilities. For two categories of attacks that are straightforward to implement, i.e., bijection and encoding, we discuss two defensive strategies, one in token and the other in embedding space. We conclude with some research questions we believe should be prioritised to enhance the defensive mechanisms of LLMs and our understanding of their safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。