提出新型多轮对话越狱代理,隐蔽性强、成功率高。
MRJ-Agent: An Effective Jailbreak Agent for Multi-Round Dialogue
- 将风险拆分至多轮对话中,结合心理策略提升攻击强度。
- 在多个测试集上达到当前最高越狱成功率,显著优于基线方法。
- 适合安全评估与对抗训练研究者,推动大模型伦理防护发展。
大型语言模型(LLMs)在知识储备和理解能力方面表现卓越,但面对越狱攻击时可能产生非法或不道德的响应。为确保其在关键应用中的负责任部署,需深入理解其安全能力和潜在漏洞。现有研究主要聚焦单轮对话越狱,忽视了多轮对话中的潜在风险——这是人类与模型交互获取信息的重要方式。尽管部分研究开始关注此问题,但通常依赖手工模板或提示工程,受限于多轮交互的复杂性,攻击效果有限。为此,本文提出一种新型多轮对话越狱代理(MRJ-Agent),强调隐蔽性,通过风险分解策略将攻击分散至多轮,并运用心理策略增强攻击力度。大量实验表明,该方法超越现有攻击手段,在多个基准测试中达到顶尖成功率。相关代码与数据集将于近期公开,供后续研究使用。
原文摘要 · Abstract (English)
Large Language Models (LLMs) demonstrate outstanding performance in their reservoir of knowledge and understanding capabilities, but they have also been shown to be prone to illegal or unethical reactions when subjected to jailbreak attacks. To ensure their responsible deployment in critical applications, it is crucial to understand the safety capabilities and vulnerabilities of LLMs. Previous works mainly focus on jailbreak in single-round dialogue, overlooking the potential jailbreak risks in multi-round dialogues, which are a vital way humans interact with and extract information from LLMs. Some studies have increasingly concentrated on the risks associated with jailbreak in multi-round dialogues. These efforts typically involve the use of manually crafted templates or prompt engineering techniques. However, due to the inherent complexity of multi-round dialogues, their jailbreak performance is limited. To solve this problem, we propose a novel multi-round dialogue jailbreaking agent, emphasizing the importance of stealthiness in identifying and mitigating potential threats to human values posed by LLMs. We propose a risk decomposition strategy that distributes risks across multiple rounds of queries and utilizes psychological strategies to enhance attack strength. Extensive experiments show that our proposed method surpasses other attack methods and achieves state-of-the-art attack success rate. We will make the corresponding code and dataset available for future research. The code will be released soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。