arXiv:2607.19351cs.AI2026-07

针对多智能体系统动态攻击,提出持续演化的双重防御机制。

OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks

论文配图:OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks
图 1 · 摘自论文原文
  • 通过双速率控制分离攻防学习速度,应对敌方策略与正常行为的双重漂移。
  • 在100轮部署中检测多数未知攻击,误报率低且无灾难性遗忘。
  • 适合安全关键场景下的开放世界多智能体系统防御。

基于大语言模型的多智能体系统(LLM-MAS)正广泛应用于安全关键领域,攻击者可通过智能体间通信注入恶意指令以传播有害行为。此类攻击具有双重动态性:攻击者会针对已部署防御不断优化策略,同时正常智能体行为随系统扩展而漂移。现有防御将部署视为封闭世界问题,一旦分布偏移超出训练覆盖范围即迅速失效。本文提出OpenEvoShield,一种面向LLM-MAS的共进化持续防御框架。其包含:异步速率控制器(M1)根据双漂移信号解耦攻击侧快速学习与正常侧慢速学习;正常边界更新器(M2)以慢速维持动态行为边界;基于经验权重固化(EWC)的策略集成(M3)实现快速适应且避免灾难性遗忘;能量基多粒度检测器(M4)融合节点、子图与图层级证据,识别新型异常攻击。在五个基准与四种多智能体拓扑上进行100轮部署实验,结果表明OpenEvoShield显著优于静态与持续基线,在保持低误报率的同时有效检测大量先前未见攻击。

原文摘要 · Abstract (English)

LLM-based multi-agent systems (LLM-MAS) are increasingly deployed in safety-critical applications, where adversaries inject malicious instructions through inter-agent communication to propagate harmful behaviors. Unlike static threats, these attacks are doubly dynamic: adversaries refine injection strategies against deployed defenses while normal-agent behavior drifts with system expansion. Existing defenses treat deployment as a closed-world problem and degrade rapidly once either distribution shifts beyond training coverage. We propose OpenEvoShield, a co-evolutionary continual defense framework for LLM-MAS. An asymmetric rate controller (M1) decouples fast attack-side and slow normal-side learning rates from dual drift signals. A normal-boundary updater (M2) maintains a dynamic behavioral boundary at the slow rate, while an EWC-regularized policy ensemble (M3) fast-adapts without catastrophic forgetting. An energy-based multi-granularity detector (M4) fuses node-, subgraph-, and graph-level evidence to classify novel attacks as out-of-distribution. Experiments over 100 deployment rounds across five benchmarks and four MAS topologies show that OpenEvoShield outperforms static and continual baselines, detecting most previously unseen attacks while keeping false positive rates low.

多智能体持续防御大模型安全动态攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。