提前模拟对话状态,拦截有害信息传播
SAIGuard: Communication-State Simulation for Proactive Defense of LLM Multi-Agent Systems

- 通过模拟通信状态预测消息影响,提前识别风险
- 在攻击成功前修复或重生成可疑消息,成功率提升
- 适合需要高安全性的多智能体协作系统
基于大语言模型的多智能体系统(MAS)通过智能体间协作完成复杂任务,但其通信驱动特性也使安全风险可在智能体间扩散并引发系统级故障。现有防御多采用执行后的被动响应策略,检测并隔离有害智能体,可能导致不可逆损害并降低协作效率。为此,我们提出一种主动防御框架——模拟感知拦截防护器(SAIGuard)。该框架对多智能体交互图进行通信状态仿真,评估传入消息对局部智能体状态及全局系统状态的影响,通过与良性通信模式的重建偏差检测潜在风险消息。不同于隔离智能体,SAIGuard在消息传播前对其进行净化或重生成。在多种拓扑结构和攻击场景下的实验表明,该方法显著降低攻击成功率,同时保持系统协作效能,优于传统被动防御。
原文摘要 · Abstract (English)
LLM-based multi-agent systems (MAS) solve complex tasks through inter-agent collaboration, but their communication-driven nature also allows security risks to spread across agents and trigger system-wide failures. Existing MAS defenses mainly follow a reactive paradigm after execution by detecting and isolating harmful agents, which may cause irreversible damage and degrade collaborative utility. To address this, we propose a proactive defense framework for MAS security, namely a Simulation-aware Interception Guard (SAIGuard). SAIGuard performs communication-state simulation over the MAS interaction graph, estimates the impact of incoming messages on local agent states and the global MAS state, and detects risky messages via reconstruction deviations from benign communication patterns. Instead of isolating agents, SAIGuard sanitizes or regenerates suspicious messages before it propagation into system. Experiments across diverse topologies and attack scenarios show that SAIGuard reduces attack success rates while maintaining MAS utility, outperforming reactive defenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。