用强化学习让提示词变得更隐蔽,突破大模型安全防线。
Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning
- 通过强化学习将攻击提示转化为形式化表达,提升隐蔽性。
- 在多个开源模型上成功绕过对齐防护,实现有效越狱。
- 结合GraphRAG结构化信息,增强攻击持续性和成功率。
大型语言模型(LLMs)展现出强大能力,但也带来新型安全挑战。例如,提示越狱攻击通过精心设计的提示诱导模型输出偏离人类价值观的内容。为揭示现有对齐方法的漏洞,我们提出PASS框架(基于语义与结构形式化的提示越狱)。该框架利用强化学习将初始攻击提示转化为形式化描述,从而增强隐蔽性并绕过现有对齐防御。生成的越狱输出被组织成GraphRAG系统,通过提取相关术语和形式化符号作为上下文输入,配合原始查询,强化后续攻击并实现更高效的越狱。我们在多个常见开源模型上进行了广泛实验,验证了该攻击的有效性。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated remarkable capabilities, yet they also introduce novel security challenges. For instance, prompt jailbreaking attacks involve adversaries crafting sophisticated prompts to elicit responses from LLMs that deviate from human values. To uncover vulnerabilities in LLM alignment methods, we propose the PASS framework (\underline{P}rompt J\underline{a}ilbreaking via \underline{S}emantic and \underline{S}tructural Formalization). Specifically, PASS employs reinforcement learning to transform initial jailbreak prompts into formalized descriptions, which enhances stealthiness and enables bypassing existing alignment defenses. The jailbreak outputs are then structured into a GraphRAG system that, by leveraging extracted relevant terms and formalized symbols as contextual input alongside the original query, strengthens subsequent attacks and facilitates more effective jailbreaks. We conducted extensive experiments on common open-source models, demonstrating the effectiveness of our attack.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。