用动态自进化框架突破大模型多轮越狱攻击瓶颈
Knowledge-Driven Multi-Turn Jailbreaking on Large Language Models
- 构建闭环规划-执行-反思系统,自主积累漏洞知识
- 在GPT-5和Claude 3.7 Sonnet上成功率显著提升
- 可抵御多种先进防御机制,适合安全研究者参考
大型语言模型面临多轮越狱攻击的严峻威胁,攻击者通过逐步引导对话获取有害输出。现有攻击方法存在连贯性差、模式僵化、难以适应模型动态状态等问题。为此,我们提出Mastermind框架,采用闭环的规划-执行-反思机制,实现对模型漏洞知识的自主构建与迭代优化。该框架采用分层规划架构,将高层目标与底层策略解耦,保障长期一致性;其知识库通过反思交互经验自动发现并优化攻击模式。基于积累的知识,系统可动态重组和调整攻击向量,显著提升攻击效果与鲁棒性。我们在GPT-5和Claude 3.7 Sonnet等前沿模型上进行实验,结果表明Mastermind显著优于现有基线,在攻击成功率和有害性评分上均有大幅提升,且对多种先进防御机制表现出强韧性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) face a significant threat from multi-turn jailbreak attacks, where adversaries progressively steer conversations to elicit harmful outputs. However, the practical effectiveness of existing attacks is undermined by several critical limitations: they struggle to maintain a coherent progression over long interactions, often losing track of what has been accomplished and what remains to be done; they rely on rigid or pre-defined patterns, and fail to adapt to the LLM's dynamic and unpredictable conversational state. To address these shortcomings, we introduce Mastermind, a multi-turn jailbreak framework that adopts a dynamic and self-improving approach. Mastermind operates in a closed loop of planning, execution, and reflection, enabling it to autonomously build and refine its knowledge of model vulnerabilities through interaction. It employs a hierarchical planning architecture that decouples high-level attack objectives from low-level tactical execution, ensuring long-term focus and coherence. This planning is guided by a knowledge repository that autonomously discovers and refines effective attack patterns by reflecting on interactive experiences. Mastermind leverages this accumulated knowledge to dynamically recombine and adapt attack vectors, dramatically improving both effectiveness and resilience. We conduct comprehensive experiments against state-of-the-art models, including GPT-5 and Claude 3.7 Sonnet. The results demonstrate that Mastermind significantly outperforms existing baselines, achieving substantially higher attack success rates and harmfulness ratings. Moreover, our framework exhibits notable resilience against multiple advanced defense mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。