arXiv:2502.11127cs.CRcs.LG2025-02ACL被引 61

用图神经网络识别大模型多智能体系统异常,提升安全防御能力

G-Safeguard: A Topology-Guided Security Lens and Treatment on LLM-based Multi-agent Systems

  • 基于多智能体对话图构建安全检测机制
  • 对抗注入攻击下性能恢复超40%
  • 适配多种大模型与大规模系统,部署灵活

基于大语言模型的多智能体系统在复杂任务中表现出色,但面临对抗攻击、信息误导和意外行为等安全隐患。为此,本文提出G-Safeguard,一种拓扑引导的安全检测与修复框架,利用图神经网络在多智能体话语图上检测异常,并通过拓扑干预实现攻击缓解。大量实验表明,该方法在多种攻击策略下均有效:可使提示注入攻击导致的性能损失恢复超过40%;对不同大模型底座和大规模多智能体系统具有高度适应性;能无缝集成至主流多智能体架构并提供安全保障。代码已开源。

原文摘要 · Abstract (English)

Large Language Model (LLM)-based Multi-agent Systems (MAS) have demonstrated remarkable capabilities in various complex tasks, ranging from collaborative problem-solving to autonomous decision-making. However, as these systems become increasingly integrated into critical applications, their vulnerability to adversarial attacks, misinformation propagation, and unintended behaviors have raised significant concerns. To address this challenge, we introduce G-Safeguard, a topology-guided security lens and treatment for robust LLM-MAS, which leverages graph neural networks to detect anomalies on the multi-agent utterance graph and employ topological intervention for attack remediation. Extensive experiments demonstrate that G-Safeguard: (I) exhibits significant effectiveness under various attack strategies, recovering over 40% of the performance for prompt injection; (II) is highly adaptable to diverse LLM backbones and large-scale MAS; (III) can seamlessly combine with mainstream MAS with security guarantees. The code is available at https://github.com/wslong20/G-safeguard.

多智能体安全防御图神经网络LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。