arXiv:2410.16155cs.CL2024-10被引 9

提出新型多智能体越狱攻击框架,揭示内存污染扩散机制

A Troublemaker with Contagious Jailbreak Makes Chaos in Honest Towns

  • 设计传染性越狱方法,让恶意指令在多智能体间快速传播
  • 在100智能体场景下攻击成功率提升52.93%
  • 揭示多智能体系统中毒性信息会逐渐消失的潜在风险

随着大语言模型在各类场景中作为智能体广泛应用,其记忆模块成为关键组件,但也易受越狱攻击。现有研究多关注单智能体或共享记忆攻击,而现实场景中记忆常为独立结构。本文提出TMCHT任务——一个大规模、多拓扑结构的文本攻击评估框架,模拟一个攻击者智能体误导整个智能体社会。我们识别出两大挑战:非完整图结构与大规模系统,归因于‘毒性消失’现象。为此,提出对抗性复制传染越狱(ARCJ)方法,通过优化检索后缀提升污染样本的可获取性,并优化复制后缀赋予其传染能力。实验表明,在线性拓扑、星型拓扑及100智能体设置下,攻击成功率分别提升23.51%、18.95%和52.93%。呼吁学界关注多智能体系统的安全性。

原文摘要 · Abstract (English)

With the development of large language models, they are widely used as agents in various fields. A key component of agents is memory, which stores vital information but is susceptible to jailbreak attacks. Existing research mainly focuses on single-agent attacks and shared memory attacks. However, real-world scenarios often involve independent memory. In this paper, we propose the Troublemaker Makes Chaos in Honest Town (TMCHT) task, a large-scale, multi-agent, multi-topology text-based attack evaluation framework. TMCHT involves one attacker agent attempting to mislead an entire society of agents. We identify two major challenges in multi-agent attacks: (1) Non-complete graph structure, (2) Large-scale systems. We attribute these challenges to a phenomenon we term toxicity disappearing. To address these issues, we propose an Adversarial Replication Contagious Jailbreak (ARCJ) method, which optimizes the retrieval suffix to make poisoned samples more easily retrieved and optimizes the replication suffix to make poisoned samples have contagious ability. We demonstrate the superiority of our approach in TMCHT, with 23.51%, 18.95%, and 52.93% improvements in line topology, star topology, and 100-agent settings. Encourage community attention to the security of multi-agent systems.

越狱攻击多智能体安全评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。