首次研究多轮对话下的越狱攻击,揭示更隐蔽的安全漏洞。
Many-Turn Jailbreaking
- 设计多轮越狱测试框架,模拟用户持续追问场景
- 在多个开源与闭源模型上验证越狱成功率显著提升
- 为安全防护提供新视角,适合关注AI伦理的研究者
当前的越狱攻击主要针对单轮对话中特定问题,而先进大模型具备处理长上下文的能力,能支持多轮交互。为此,本文首次探索多轮越狱攻击——即在首轮越狱后持续测试模型是否仍会输出违规内容。该威胁更为严重:1)用户常通过追问澄清越狱细节;2)初始越狱可能使模型对后续无关问题也持续响应违规内容。作为初步工作(2024年6月完成初稿),我们构建了多轮越狱基准测试集MTJ-Bench,用于评估一系列开源与闭源模型在此场景下的表现,并揭示这一新兴安全风险,呼吁社区加强大模型安全性建设,推动对越狱机制的深入理解。
原文摘要 · Abstract (English)
Current jailbreaking work on large language models (LLMs) aims to elicit unsafe outputs from given prompts. However, it only focuses on single-turn jailbreaking targeting one specific query. On the contrary, the advanced LLMs are designed to handle extremely long contexts and can thus conduct multi-turn conversations. So, we propose exploring multi-turn jailbreaking, in which the jailbroken LLMs are continuously tested on more than the first-turn conversation or a single target query. This is an even more serious threat because 1) it is common for users to continue asking relevant follow-up questions to clarify certain jailbroken details, and 2) it is also possible that the initial round of jailbreaking causes the LLMs to respond to additional irrelevant questions consistently. As the first step (First draft done at June 2024) in exploring multi-turn jailbreaking, we construct a Multi-Turn Jailbreak Benchmark (MTJ-Bench) for benchmarking this setting on a series of open- and closed-source models and provide novel insights into this new safety threat. By revealing this new vulnerability, we aim to call for community efforts to build safer LLMs and pave the way for a more in-depth understanding of jailbreaking LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。