研究多智能体辩论系统如何被结构化越狱攻击,漏洞比单模型高52个百分点。
Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate
- 设计叙事包裹、角色升级等四步攻击框架,利用多智能体互动机制
- 攻击使有害内容比例从28.14%升至80.34%,最高成功率80%
- 揭示多智能体系统固有安全缺陷,适合关注AI安全的研究者
多智能体辩论(MAD)通过大型语言模型(LLMs)间的协作互动提升复杂任务推理能力。然而,其迭代对话与角色扮演特性带来的安全风险,尤其是对越狱攻击诱发有害内容的脆弱性,仍严重缺乏研究。本文系统分析了基于GPT-4o、GPT-4、GPT-3.5-turbo和DeepSeek等主流商业LLM构建的四种典型MAD框架的安全性,未修改内部代理。提出一种新型结构化提示重写框架,专门利用叙事封装、角色驱动升级、迭代优化和修辞模糊等手段攻击MAD动态。大量实验表明,MAD系统比单智能体设置更易受攻击。关键的是,所提方法显著放大脆弱性,使平均有害性从28.14%提升至80.34%,在特定场景下攻击成功率高达80%。研究揭示了MAD架构的内在漏洞,强调在实际部署前亟需建立专用鲁棒防御机制。
原文摘要 · Abstract (English)
Multi-Agent Debate (MAD), leveraging collaborative interactions among Large Language Models (LLMs), aim to enhance reasoning capabilities in complex tasks. However, the security implications of their iterative dialogues and role-playing characteristics, particularly susceptibility to jailbreak attacks eliciting harmful content, remain critically underexplored. This paper systematically investigates the jailbreak vulnerabilities of four prominent MAD frameworks built upon leading commercial LLMs (GPT-4o, GPT-4, GPT-3.5-turbo, and DeepSeek) without compromising internal agents. We introduce a novel structured prompt-rewriting framework specifically designed to exploit MAD dynamics via narrative encapsulation, role-driven escalation, iterative refinement, and rhetorical obfuscation. Our extensive experiments demonstrate that MAD systems are inherently more vulnerable than single-agent setups. Crucially, our proposed attack methodology significantly amplifies this fragility, increasing average harmfulness from 28.14% to 80.34% and achieving attack success rates as high as 80% in certain scenarios. These findings reveal intrinsic vulnerabilities in MAD architectures and underscore the urgent need for robust, specialized defenses prior to real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。