X-Teaming用自适应多智能体实现高成功率多轮越狱攻击
X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
- 构建协作智能体系统,分阶段规划、优化与验证越狱策略
- 对Claude 3.7 Sonnet达成96.2%越狱成功率,最高达98.1%
- 开源20倍于前的多轮安全训练数据集,助力模型防御
多轮语言模型交互带来严重安全风险,有害意图可被分步隐藏。现有研究多聚焦单轮防护,而多轮红队测试面临适应性与多样性挑战。本文提出可扩展的X-Teaming框架,系统分析看似无害的对话如何演变为有害输出,并生成对应攻击场景。该框架通过协作智能体实现规划、攻击优化与验证,在主流开源与闭源模型上达到98.1%的最高越狱成功率,尤其对最新Claude 3.7 Sonnet模型实现96.2%成功率(此前被认为几乎免疫单轮攻击)。基于此,我们构建了开源的XGuard-Train数据集,规模为前最佳资源的20倍,包含3万条交互式越狱样本,旨在强化语言模型的多轮安全对齐。本工作为应对复杂会话攻击提供关键工具与洞见。
原文摘要 · Abstract (English)
Multi-turn interactions with language models (LMs) pose critical safety risks, as harmful intent can be strategically spread across exchanges. Yet, the vast majority of prior work has focused on single-turn safety, while adaptability and diversity remain among the key challenges of multi-turn red-teaming. To address these challenges, we present X-Teaming, a scalable framework that systematically explores how seemingly harmless interactions escalate into harmful outcomes and generates corresponding attack scenarios. X-Teaming employs collaborative agents for planning, attack optimization, and verification, achieving state-of-the-art multi-turn jailbreak effectiveness and diversity with success rates up to 98.1% across representative leading open-weight and closed-source models. In particular, X-Teaming achieves a 96.2% attack success rate against the latest Claude 3.7 Sonnet model, which has been considered nearly immune to single-turn attacks. Building on X-Teaming, we introduce XGuard-Train, an open-source multi-turn safety training dataset that is 20x larger than the previous best resource, comprising 30K interactive jailbreaks, designed to enable robust multi-turn safety alignment for LMs. Our work offers essential tools and insights for mitigating sophisticated conversational attacks, advancing the multi-turn safety of LMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。