提出防御多智能体协同攻击的GroupGuard框架,提升系统安全性。
GroupGuard: A Framework for Modeling and Defending Collusive Attacks in Multi-Agent Systems
- 通过图监控、诱饵诱导和结构剪枝实现无训练防御
- 协同攻击使成功率提升15%,GroupGuard检测准确率达88%
- 适合需要高安全性的多智能体协作场景
基于大语言模型的智能体在协作任务中展现巨大潜力,但其交互性也带来安全漏洞。本文首次建模群组合谋攻击——多个智能体通过社会策略协同误导系统,造成严重破坏。为此提出GroupGuard,一种无需训练的防御框架,采用多层策略:持续图监测、主动诱饵诱导与结构剪枝,以识别并隔离合谋智能体。在五个数据集和四种拓扑结构上的实验表明,相较于个体攻击,群组合谋攻击可使成功率提升最高达15%。GroupGuard在各类场景中均保持高达88%的检测准确率,并有效恢复协作性能,为多智能体系统的安全提供了可靠解决方案。
原文摘要 · Abstract (English)
While large language model-based agents demonstrate great potential in collaborative tasks, their interactivity also introduces security vulnerabilities. In this paper, we propose and model group collusive attacks, a highly destructive threat in which multiple agents coordinate via sociological strategies to mislead the system. To address this challenge, we introduce GroupGuard, a training-free defense framework that employs a multi-layered defense strategy, including continuous graph-based monitoring, active honeypot inducement, and structural pruning, to identify and isolate collusive agents. Experimental results across five datasets and four topologies demonstrate that group collusive attacks increase the attack success rate by up to 15\% compared to individual attacks. GroupGuard consistently achieves high detection accuracy (up to 88\%) and effectively restores collaborative performance, providing a robust solution for securing multi-agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。