提出防御大模型多智能体系统传播攻击的新框架,能精准追踪并修复污染路径。
PropGuard: Safeguarding LLM-MAS via Propagation-Aware Exploration and Remediation

- 构建时空双视角图,结合响应风险与状态证据追踪攻击传播。
- 在五种攻击设置下降低攻击成功率,保持任务完成率超90%。
- 适合需要高安全性的协作式AI系统开发者使用。
基于大语言模型的多智能体系统(LLM-MAS)通过角色分工、工具调用、记忆存储和协同推理解决复杂任务,但消息、工具或记忆中的恶意指令可在智能体间及多轮交互中传播,导致系统级失效。现有防御依赖局部过滤或图异常检测,难以追踪细粒度传播路径,且修复时易破坏正常协作。本文提出PropGuard,一种感知传播的防护框架。它构建双视角时空图,融合以响应为中心的风险估计与全状态证据保留;基于风险先验,采用GE-GRPO训练的检查员逐步探索全状态图,恢复紧凑的可疑传播子图;再通过子图感知诊断验证危害传播,并实施源导向修复,纠正上游污染并重播受影响的下游交互。在四种通信架构与五类攻击场景下的实验表明,PropGuard持续降低攻击成功率,同时维持高任务级防御成功率,实现良好有效性和效率平衡。
原文摘要 · Abstract (English)
LLM-based multi-agent systems (LLM-MAS) have become a promising paradigm for solving complex tasks through role specialization, tool use, memory, and collaborative reasoning. However, these interactions create new security risks that malicious instructions injected through messages, tools, or memories can propagate across agents and rounds, causing system-level compromise. Existing defenses largely rely on local filtering or graph-based anomaly detection, but they often fail to trace fine-grained propagation paths or remediate contaminated states without disrupting benign collaboration. We propose PropGuard, a propagation-aware framework for safeguarding LLM-MAS. PropGuard constructs a dual-view spatio-temporal graph that combines response-centric risk estimation with full-state evidence preservation. Guided by these risk priors, a GE-GRPO trained inspector sequentially explores the full-state graph to recover compact suspicious propagation subgraphs. PropGuard then verifies harmful propagation through subgraph-aware diagnosis and applies source-guided remediation to correct upstream contamination and replay affected downstream interactions. Experiments across four communication architectures and five attack settings demonstrate that PropGuard consistently lowers attack success while maintaining high task-level defense success, achieving a favorable effectiveness--efficiency trade-off.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。